Search and filter through extracted claims from AI researchers.
Showing 181-200 of 494 claims of type "critique"
Claims about a golden age of science from Astra are a leap of faith and overgeneralization from formal to difficult-to-formalize problems
The AI community commits a logical fallacy by assuming that success in one cognitive domain (like math) means imminent success in all cognitive domains
Existing medical image fusion methods lack deep understanding of diagnostic intents and pathological structures by applying uniform fusion rules globally
Current multimodal RAG systems struggle with complex multi-hop reasoning because they focus on instance-level matching and fail to capture relationships across modalities and documents
High-fidelity 3D generation's reliance on scaling model capacity and data incurs prohibitive computational costs and overlooks rich priors in discriminative 3D foundation models
On-policy self-distillation (OPSD) remains brittle in practice and requires substantial engineering effort to work reliably
SWE-bench-like benchmarks suffer from systematic misalignment due to the complexity of PR-Issue pairing in large repositories
Existing post-training quantization methods for Vision Transformers use uniform bit-widths and overlook heterogeneous sensitivity to quantization, leading to inefficient precision allocation
Multimodal on-policy distillation's next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects, making it difficult to isolate visual evidence
Existing agentic visual reasoning systems fail to optimize for Mode Adaptiveness and Tool Effect, leading to inefficient tool use
Current safety alignment efforts inadvertently alter models' representations of mindedness in other entities alongside human beliefs and values
Vision-language models as judges of computer-using agent trajectories have not been systematically evaluated for reliability
System prompts are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in AI system deployment
LLM-as-a-judge for GenUI evaluation is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users perceive interfaces
Most existing 3D-QA methods rely on costly 3D-specific training or fine-tuning with annotations, limiting their scalability and real-world applicability
On-policy distillation effectiveness depends on teacher consistency, where the OPD supervision model should have generated the SFT demonstrations, but this condition is frequently violated in practice
Existing visual sampling methods for long videos either produce redundant frame selection with insufficient temporal coverage or use fixed strategies regardless of query type
Financial sentiment analysis reduces rich, multi-dimensional news articles to a single polarity score, missing orthogonal information dimensions
Most existing algorithmic recourse approaches focus only on flipping predictions without accounting for genuine improvement in qualifications versus strategic gaming
GenAI in academic writing raises concerns about reinforcing dominant language norms and marginalizing diverse Englishes
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.