Search and filter through extracted claims from AI researchers.
Showing 61-80 of 435 claims in topic "multimodal"
"Existing pipelines, however, assume that material parameters are known or manually specified, limiting their applicability when these parameters must be inferred from observed object dynamics."
"We propose KnockGS, an interaction-response PhysicalGS framework that estimates the elasticity and density scales of a 3D Gaussian object from its dynamics under a known applied force."
"Across five held-out material targets, our method recovers the scales substantially more accurately than response retrieval, global regression, or a fixed default material"
"Interaction response therefore carries enough information to calibrate material scales in physically grounded 3D Gaussian representations."
"Our study is a first step toward interactive PhysicalGS systems that calibrate a Gaussian asset whose rendered appearance and simulated response are consistent."
"Although this setting better reflects natural storytelling, later prompts may omit important identity-related semantics, making character consistency more difficult to maintain."
"Experiments on FreeStoryBench show that Sidecar consistently improves prompt-image alignment and character consistency across multiple SDXL- and FLUX-based baselines, with negligible computational overhead."
"it is impossible to project features into a common space that is adaptive for all the tasks"
"The experimental results demonstrate that the proposed deformation head shows good results on other detection methods"
The TADP method achieves 80.91% car mAP on KITTI dataset, surpassing many state-of-the-art methods
"The experimental results on the KITTI dataset demonstrate that the car mAP is 80.91%, surpassing many state-of-the-art methods on the KITTI benchmark"
"the common paradigm of directly concatenating multispectral sequences can exhibit anomalous performance degradation, primarily due to pronounced heterogeneity and the resulting multimodal imbalance across modalities"
"MM-Spectrum introduces an explicit modality-aware routing mechanism that exposes spectral identity to the router in addition to token content representations"
"it incorporates shared and interaction experts, together with heterogeneous expert capacities, to extract multispectral modality-unique and cross-modal synergistic information while suppressing noise-induced interference"
"Across full-modality, bimodal, and missing-modality settings on molecular structural elucidation, MM-Spectrum achieves consistent and substantial improvements, supported by ablation studies and interpretability analyses"
"Although computer vision has advanced significantly, most state-of-the-art models perform well on standard benchmarks but often yield suboptimal results in specialized medical imaging tasks due to the high level of noise present in the data."
"Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors."
"Recent video generation models are increasingly framed as world models."
"a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action"
"existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution"
"we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.