Search and filter through extracted claims from AI researchers.
Showing 161-180 of 435 claims in topic "multimodal"
"We believe that jointly advancing omni-modal representation learning and omni-interactive querying paves the way toward universal embedders."
"A DM is trained to learn the mapping between the parameter space and the observable space."
"positive frames account for under 3% of a sequence, form short contiguous segments, and are poorly handled by off-the-shelf ultrasound and vision foundation models"
A new large-scale dataset of 115K scenes is the first hybrid dataset for image-to-scene generation
"we construct a new large-scale dataset of 115K scenes. To the best of our knowledge, it is the first hybrid dataset for image-to-scene generation."
"Detecting the fetal abdominal circumference standard plane in low-cost obstetric blind sweeps is a highly imbalanced frame-classification problem: positive frames account for under 3% of a sequence"
"On the ACOUSLIC-AI benchmark, AnatoProto reaches a test F1 of 67.72, outperforming the strongest foundation-model baseline (FetalCLIP + PRS, F1 = 54.52) by +13.20 F1 and the strongest video temporal-action-detection baseline (TriDet + PRS) by +15.76 F1"
"A synergy study, backed by embedding geometry and paired-bootstrap confidence intervals, shows that the prototype loss and anatomy-weighted pooling are not additive: applied alone the prototype loss reduces recall by 12 points, but combined with anatomy-weighted pooling it increases recall by 6.5 points"
"anatomy-weighted spatial pooling that uses nnU-Net abdominal-region probabilities as a spatial prior to reweight BiomedCLIP patch tokens, so frozen semantic features are aggregated onto anatomically meaningful regions"
"On-policy self-distillation (OPSD) has recently emerged as an effective post-training paradigm that improves policy optimization through dense token-level supervision from a privileged self-teacher."
OPSD remains largely underexplored for Video Large Language Models despite its promise
"Despite its promise, OPSD remains largely underexplored for Video Large Language Models (Video-LLMs)."
"long videos contain substantial temporal redundancy, and only a small subset of frames provides the evidence necessary to answer a question."
Video-OPSD achieves performance comparable to GRPO while requiring substantially less training time
"achieves performance comparable to GRPO while requiring substantially less training time"
"This focused visual input enables the teacher to provide more informative supervision."
"Existing methods based on video diffusion model (VDM) commonly rely on incomplete conditioning signals such as sparse point clouds or 2D panoramas, leading to stochastic hallucinations, long-term drifts and suboptimal 3D consistency."
"Extensive experiments on both synthetic and real-world datasets show that SpatialCrafter outperforms state-of-the-art methods, mitigates long-term drift, and remains robust and consistent under rapid camera motion and extreme viewpoint changes."
"we propose a Point-anchored Sparse Structure~(PaSS) Flow module that predicts a spatially aligned and geometrically consistent 3D proxy"
"Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation."
"we find that utility is strongly task-conditioned. Gains concentrate in visual cue injection, grounding, and counterfactual state realization"
"intermediates requiring symbol-sensitive construction or structural extrapolation are substantially less reliable"
"our consolidated Qwen pipeline improves the mean task score from 0.343 to 0.445 ($+10.2$ points; $+29.7\%$ relative)"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.