Search and filter through extracted claims from AI researchers.
Showing 261-280 of 435 claims in topic "multimodal"
HyperClaim framework enables better localized authenticity detection by constructing sparse heterogeneous hypergraphs over query tokens, evidence tokens, and sampled frames
Foundation models introduce transferable prior knowledge that offers new ways to address hand-object interaction modeling challenges beyond task-specific data and models
The literature on foundation models for hand-object interaction remains fragmented, with studies typically describing methods simply as 'using large models' without systematic characterization
Most segmentation algorithms lack the generalisation capacity required for large-scale flood monitoring application, while annotated flood data are scarce and unevenly distributed
Existing single-image head avatar reconstruction approaches struggle to preserve 3D consistency under unseen viewpoints
Dynamic 3D head avatars can be rendered in real time using deformable 3D Gaussian Splatting with binding templates
LLM pipelines can achieve strong correlation (0.867) with expert ratings for automated depression assessment in clinical trials
Existing automated depression detection solutions provide limited support for clinical trials with structured interviews
Model scale shows weak correlation with robustness to multi-attribute biases in vision-language models, with Spearman correlation dropping from 0.68 on ImageNet to only 0.05 on multi-attribute bias benchmarks
Training data quality is more important than model scale for VLM robustness, with curated datasets yielding up to 25% improvements over uncurated alternatives at comparable scale
VLM robustness to spurious correlations remains poorly understood at scale despite CLIP-like models being foundational to multimodal systems
Collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly for embodied AI
Augmenting 400 real trajectories with 400 synthetically generated egocentric manipulation videos improves visual fidelity, geometric stability, and action alignment in long egocentric rollouts
Integrating Sentinel-1 SAR and Sentinel-2 multispectral satellite data with street-level imagery from Mapillary can overcome cloud-induced temporal gaps and provide more accurate parcel-level agricultural monitoring
An automated pipeline can successfully filter and curate crowdsourced street-level imagery at scale, producing 46,050 analysis-ready annotated images from over 900,000 initial images
Camera motion in video generation is primarily established during high-noise stages of the diffusion process, where coarse spatiotemporal structures are formed
Video re-shooting can be achieved without explicit 3D priors or paired training data by using text-driven semantic viewpoint specification and self-supervised learning of camera dynamics
Existing disaster response datasets like Incidents1M and CrisisMMD suffer from either complete lack of text or severe text-image semantic misalignment
High-fidelity textual descriptions for vision-only datasets can be successfully generated and validated at scale using a combination of dense and MoE language models plus LLM-as-a-Judge validation
An image-blind LLM-as-a-Judge validation approach can ensure generated captions provide reliable semantic anchoring for Data-Free Knowledge Distillation by intentionally obscuring the original image
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.