Search and filter through extracted claims from AI researchers.
Showing 201-220 of 494 claims of type "critique"
Making evidence available does not ensure that complementary cues across moments are integrated for answering in video understanding
Existing evaluations of graph learning rely on incompatible splits and assumptions that make conclusions about same-graph cross-task transfer unreliable
Existing multi-agent systems typically treat communication topology as a fixed design choice or an offline optimization target, which is a limitation.
AI technologies reproduce dominant language ideologies at multiple levels including training data, design protocols, evaluation benchmarks, user feedback and public commentary.
Dominant multimodal benchmarks in pathology mainly score final answers but provide limited insight into whether models understand multiscale visual content needed for pathology reasoning
Trajectory-level rewards in on-policy distillation cannot determine whether a failed answer arose from perception or subsequent reasoning
Existing training-free visual token pruning methods suffer from prematurely discarding tokens essential for deep-layer reasoning due to reliance on static, instantaneous heuristics
Computer-use agent benchmark scores are commonly produced by brittle scripted oracles that can produce unreliable results
Multimodal agent evaluation that reduces to final-answer accuracy cannot distinguish whether correct answers came from grounded evidence, language priors, or accidental error cancellation
There is a lack of controllable and attributable methods for analyzing how language models resolve conflicts between competing specifications
Existing gradient-based explainability methods struggle to provide global insights into what specifically drives similarity in regions of an embedding space
Extending order-optimal convergence guarantees to neural critics in average-reward CMDPs has remained an open problem due to a fundamental bias-cost trade-off
The literature on foundation models for hand-object interaction remains fragmented, with studies typically describing methods simply as 'using large models' without systematic characterization
Computer-use agents often fail on transient GUI events because expensive autoregressive decoding is on the decision-time critical path
Most segmentation algorithms lack the generalisation capacity required for large-scale flood monitoring application, while annotated flood data are scarce and unevenly distributed
Existing post-training quantization methods are poorly suited to World Action Models because they rely on open-loop objectives, homogeneous model assumptions, and calibration distributions that do not reflect deployment
Existing visual-latent reasoning methods fail to fully internalize the abstract reasoning process induced by multimodal Chain-of-Thought
Existing multimodal long-term memory agents lack mechanisms to diagnose retrieval failures and adapt search strategies
Existing single-image head avatar reconstruction approaches struggle to preserve 3D consistency under unseen viewpoints
Existing early-exit gates for diffusion language models fire prematurely on long chain-of-thought outputs whose answers stabilize only near the end
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.