Search and filter through extracted claims from AI researchers.
Showing 101-120 of 435 claims in topic "multimodal"
"On two further datasets with severe distribution shift, every model including HALO collapses zero-shot"
"aligns this IMU encoder with text embeddings via synonym-aware soft contrastive learning, enabling open-set recognition via cosine-similarity retrieval without per-dataset classifiers"
"Research into event-based object classification methods are hindered by the lack of high-quality vision datasets to use."
"This work simultaneously provides four datasets with rich details for future experiments to use and validates the output of the ANTShapes dataset simulation tool as being suitable for its purpose."
"This approach has several practical flaws. The size, weight and power consumption of the device could prohibit deployment at the extreme edge or in covert sensing environments. Besides this, there are security concerns inherent in cloud-based or other off-device computation approaches due to the requirement of sending and receiving potentially sensitive data. Furthermore, this transmission of data introduces latency and requires consistent connectivity to the cloud infrastructure to function."
"Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images."
"Our experiments reveal a striking performance gap between even the best VLMs and human"
Most VLMs struggle to accurately perceive visual text, resulting in frequent correction errors
"most models struggle to accurately perceive the visual text, resulting in frequent correction errors"
"which requires a profound understanding of the interplay between visual text and its surrounding visual context"
Video foundation models are beginning to change film and video production
"Recently, video foundation models are beginning to change film and video production"
"Magpie provides a system-level implementation path for applying generative models to real-time game rendering. It preserves gameplay designability and reproducibility, and reduces the dependence of early game prototypes on complete visual assets."
Generative models can reduce the dependence of early game prototypes on complete visual assets
"reduces the dependence of early game prototypes on complete visual assets"
"existing evaluations often allow shortcuts based on transcripts or single-modality solutions, obscuring whether models genuinely ground predictions in speech"
"strong text-only LLMs exceed 90% accuracy in consistent cases but drop to 33-48% in conflict cases"
"Direct AudioLLMs provide only partial grounding, still selecting the transcript-biased trap in roughly 30-40% of conflict cases"
"explicit acoustic evidence aggregation provides a more controllable interface for diagnosing and improving speech-grounded reasoning"
"These results identify transcript-based shortcuts as an important failure mode in spoken dialogue understanding"
"A Swin V2 encoder pretrained on 10,444 public 3D CT volumes using a DINOv2-style objective was adapted to T2-weighted MRI through four cumulative configurations"
"SWIFTe reduced total parameters by 70.1% (from 72.8M to 21.8M) and increased tumor detection rate from 89.9% to 93.9%"
"removing tumor-aware augmentation reduced detection from 93.9% to 89.9% but increased surface DSC from 0.61 to 0.64, demonstrating a detection-boundary-agreement trade-off"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.