HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 21-40 of 278 claims in topic "benchmarks"

benchmarks
fact
Bullish
academic

R2M-Bench's Overall NMR metric correlates with human consistency judgments at Spearman's ρ=0.547

"Overall NMR correlates with human consistency judgments at Spearman's $ρ=0.547$ (95\% CI $[0.45,0.63]$)."
Computer Vision
8/30/2026
Confidence: 90%Source
Previous
1314
Page 2 of 14
benchmarks
fact
Bullish
academic

Relative calibration substantially reduces the slow-motion shortcut in video world model evaluation, with within-model correlation magnitude of 0.072 compared to 0.207 for raw revisit similarity

"Its within-model correlation magnitude with generated motion is $0.072$, compared with $0.207$ for raw revisit similarity, indicating that relative calibration substantially reduces the slow-motion shortcut."
Computer Vision
8/30/2026
Confidence: 85%Source
benchmarks
fact
Neutral
academic

DreamX-World-Memo achieves the highest Overall NMR among evaluated video models

"DreamX-World-Memo achieves the highest Overall NMR among the evaluated video models."
Computer Vision
8/30/2026
Confidence: 90%Source
benchmarks
opinion
Bullish
academic

Same-rollout relative calibration is a practical way to distinguish revisit-specific consistency from generic temporal stability in video world models

"these results support same-rollout relative calibration as a practical way to distinguish revisit-specific consistency from generic temporal stability."
Computer Vision
8/30/2026
Confidence: 80%Source
benchmarks
fact
Neutral
academic

Industrial sites contain large volumes of read-only telemetry, but few benchmarks specify how to compile these records into executable multi-turn agent tasks

"Industrial sites contain large volumes of read-only telemetry, but few benchmarks specify how to compile these records into executable multi-turn agent tasks."
Computation and Language
8/30/2026
Confidence: 85%Source
benchmarks
fact
Bullish
academic

BTS-AgentBench provides a telemetry-to-episode construction method that successfully produces 532 rows with multiple advanced features including clarification, goal revision, and evidence attribution

"The 532-row release adds clarification, goal revision, timestamp policy, quality-gated reporting, and evidence attribution while preserving the source computation and split."
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Bullish
academic

The BTS-AgentBench construction is fully reproducible, with two independent builds matching all logical exports and reproducing the exact same train/dev/test artifact

"Two independent raw-to-episode builds match all 11 logical tool-store exports and reproduce the released 356/87/89 train/dev/test artifact exactly."
Computation and Language
8/30/2026
Confidence: 95%Source
benchmarks
fact
Bullish
academic

GPT-5.5 achieves perfect completion on the XAI4HEAT held-out test split when evaluated using the shared construction path

"on its 41-row held-out test split, the controller completes 0 rows and the retained GPT-5.5 execution completes all 41."
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Bearish
academic

Document-to-document LLM translation in a single pass frequently suffers from structural misalignment including sentence omissions or hallucinations

"document-to-document generation in a single pass frequently suffers from structural misalignment, manifesting as sentence omissions or hallucinations that violate the core requirement of source-target correspondence"
Computation and Language
8/30/2026
Confidence: 85%Source
benchmarks
fact
Bullish
academic

StarPO framework significantly enhances translation quality and structural integrity in document-level machine translation

"Experimental results across news and literary domains demonstrate that StarPO significantly enhances translation quality and structural integrity"
Computation and Language
8/30/2026
Confidence: 80%Source
benchmarks
fact
Bullish
academic

Compact models using StarPO can surpass GPT-4o performance while maintaining superior token efficiency

"StarPO allows compact models to surpass the performance of massive proprietary systems like GPT-4o while maintaining superior token efficiency"
Computation and Language
8/30/2026
Confidence: 75%Source
benchmarks
fact
Bearish
academic

Ancient Chinese artifact text recognition remains fundamentally unsolved despite advances in Vision-Language Models and OCR-specialist models

"Extensive experiments on Ancient-Bench covering general Vision-Language Models (VLMs) and OCR-specialist models reveal that ancient Chinese artifact text recognition remains fundamentally unsolved, with persistent challenges in variant characters, specialized symbols, and hallucination."
Computer Vision
8/30/2026
Confidence: 90%Source
benchmarks
critique
Neutral
academic

Existing benchmarks for ancient Chinese text recognition suffer from fragmentation in temporal coverage, medium diversity, and script type completeness

"However, existing benchmarks suffer from ''fragmentation'', manifested in limited temporal coverage, limited medium diversity, and incomplete script types."
Computer Vision
8/30/2026
Confidence: 85%Source
benchmarks
fact
Bearish
academic

Current models face persistent challenges with variant characters, specialized symbols, and hallucination when processing ancient Chinese texts

"persistent challenges in variant characters, specialized symbols, and hallucination"
Computer Vision
8/30/2026
Confidence: 90%Source
benchmarks
fact
Neutral
academic

LLM agents are increasingly being applied to anomaly detection and root-cause analysis in time-series observations from real-world systems

"LLM agents are increasingly applied to anomaly detection and root-cause analysis in time-series observations collected from real-world systems"
Machine Learning
8/30/2026
Confidence: 90%Source
benchmarks
critique
Neutral
academic

LLM agent performance on anomaly detection and root-cause analysis tasks has not been systematically evaluated under controlled conditions

"their performance on these tasks has not been systematically evaluated under controlled conditions"
Machine Learning
8/30/2026
Confidence: 90%Source
benchmarks
fact
Neutral
academic

LLM agents benefit substantially from domain context when analyzing time-series observations from dynamical systems

"Our results show that agents benefit substantially from domain context"
Machine Learning
8/30/2026
Confidence: 85%Source
benchmarks
fact
Neutral
academic

LLM agents explore time-series data primarily through numerical console output rather than visualizations

"explore data primarily through numerical console output rather than visualizations"
Machine Learning
8/30/2026
Confidence: 85%Source
benchmarks
fact
Bearish
academic

LLM agents generally perform worse when required to produce a Python script for root-cause prediction than when submitting predictions directly

"agents generally perform worse when required to produce a Python script that maps each time-series sample to a predicted root-cause label than when they submit predictions directly"
Machine Learning
8/30/2026
Confidence: 85%Source
benchmarks
fact
Neutral
academic

Reliable assessment of causal research designs in social sciences has relied entirely on manual expert analysis until now

"Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis."
Computation and Language
8/30/2026
Confidence: 90%Source
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,949 pending.