Search and filter through extracted claims from AI researchers.
Showing 21-40 of 278 claims in topic "benchmarks"
R2M-Bench's Overall NMR metric correlates with human consistency judgments at Spearman's ρ=0.547
"Overall NMR correlates with human consistency judgments at Spearman's $ρ=0.547$ (95\% CI $[0.45,0.63]$)."
"Its within-model correlation magnitude with generated motion is $0.072$, compared with $0.207$ for raw revisit similarity, indicating that relative calibration substantially reduces the slow-motion shortcut."
DreamX-World-Memo achieves the highest Overall NMR among evaluated video models
"DreamX-World-Memo achieves the highest Overall NMR among the evaluated video models."
"these results support same-rollout relative calibration as a practical way to distinguish revisit-specific consistency from generic temporal stability."
"Industrial sites contain large volumes of read-only telemetry, but few benchmarks specify how to compile these records into executable multi-turn agent tasks."
"The 532-row release adds clarification, goal revision, timestamp policy, quality-gated reporting, and evidence attribution while preserving the source computation and split."
"Two independent raw-to-episode builds match all 11 logical tool-store exports and reproduce the released 356/87/89 train/dev/test artifact exactly."
"on its 41-row held-out test split, the controller completes 0 rows and the retained GPT-5.5 execution completes all 41."
"document-to-document generation in a single pass frequently suffers from structural misalignment, manifesting as sentence omissions or hallucinations that violate the core requirement of source-target correspondence"
"Experimental results across news and literary domains demonstrate that StarPO significantly enhances translation quality and structural integrity"
"StarPO allows compact models to surpass the performance of massive proprietary systems like GPT-4o while maintaining superior token efficiency"
"Extensive experiments on Ancient-Bench covering general Vision-Language Models (VLMs) and OCR-specialist models reveal that ancient Chinese artifact text recognition remains fundamentally unsolved, with persistent challenges in variant characters, specialized symbols, and hallucination."
"However, existing benchmarks suffer from ''fragmentation'', manifested in limited temporal coverage, limited medium diversity, and incomplete script types."
"persistent challenges in variant characters, specialized symbols, and hallucination"
"LLM agents are increasingly applied to anomaly detection and root-cause analysis in time-series observations collected from real-world systems"
"their performance on these tasks has not been systematically evaluated under controlled conditions"
"Our results show that agents benefit substantially from domain context"
"explore data primarily through numerical console output rather than visualizations"
"agents generally perform worse when required to produce a Python script that maps each time-series sample to a predicted root-cause label than when they submit predictions directly"
"Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.