Search and filter through extracted claims from AI researchers.
Showing 241-260 of 435 claims in topic "multimodal"
Most existing 3D-QA methods rely on costly 3D-specific training or fine-tuning with annotations, limiting their scalability and real-world applicability
Existing visual sampling methods for long videos either produce redundant frame selection with insufficient temporal coverage or use fixed strategies regardless of query type
VisualRouter's query-grounded visual sampling framework that classifies queries and applies corresponding sampling strategies can improve long video understanding
EndoCLIP foundation model trained on 125,756 lesion-level image-text pairs from routine colonoscopy reports outperforms general-purpose and biomedical vision-language encoders across multiple clinical tasks
EndoCLIP's linear probe approaches expert reader performance on benign-versus-malignant classification in blinded study with 12 endoscopists
Recovering finding-to-frame correspondence from routine medical reports can transform routine documentation into foundation-model training data
RefCaptioner achieves best overall performance among open-source models for multi-reference image-grounded video captioning
Combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO can jointly improve reference selection, phrase-level binding, and cross-reference consistency while preserving general video-captioning ability
Organizing selected frames into query-relevant cross-frame evidence before generation improves long-video understanding
Making evidence available does not ensure that complementary cues across moments are integrated for answering in video understanding
Prior work in optical chemical structure recognition (OCSR) for single molecules is quite mature, while Markush structure parsing remains a challenging task.
OCSRGlyph achieves state-of-the-art performance in optical chemical structure recognition by carefully considering stereochemistry.
MarkushGlyph improves upon prior Markush structure parsing systems by reading the entire structure as an image with a vision-language model, rather than using multiple stages to separately process visual and text input.
Trajectory-level rewards in on-policy distillation cannot determine whether a failed answer arose from perception or subsequent reasoning
Perception-Correction Distillation identifies correctable perception failures using downstream failure and teacher-student disagreement as complementary witnesses
Existing training-free visual token pruning methods suffer from prematurely discarding tokens essential for deep-layer reasoning due to reliance on static, instantaneous heuristics
Token importance in MLLMs evolves dynamically across layers rather than remaining fixed, requiring temporal trajectory modeling instead of snapshot decisions
Trend-aware Pruning can selectively reactivate late-blooming tokens that are initially undervalued but exhibit rising semantic importance
Speech deepfake detection systems can improve generalization through training recipe enhancements (multi-source data, attack-balanced sampling, audio augmentation) rather than additional architectural complexity
Pairwise graph formulations are insufficient to capture multi-way cross-modal dependencies in video misinformation detection, whereas hypergraphs offer a suitable representation
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.