HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 81-100 of 278 claims in topic "benchmarks"

benchmarks
fact
Bullish
independent

Superforecasting transformed from an obscure academic subfield to a multibillion dollar industry through prediction markets

Scott Alexander
8/4/2026
Confidence: 90%Source
benchmarks
Previous
146
fact
Bullish
independent

AI superforecasters have come close to the accuracy of top humans, with performance rising rapidly

Scott Alexander
8/4/2026
Confidence: 85%Source
benchmarks
prediction
Neutral
independent

In a year or two, either AIs will plateau at or slightly above human forecasting levels, or they will achieve ultraforecasting capabilities far beyond top humans

Scott Alexander
8/4/2026
Confidence: 70%Source
benchmarks
opinion
Bearish
independent

Daniel Reeves argues that humans have already come close to some fundamental limit on the predictability of world events, implying AI forecasting will plateau

Scott Alexander
8/4/2026
Confidence: 75%Source
benchmarks
fact
Neutral
lab researcher

All benchmarks are trending toward saturation

Cristobal Valenzuela
8/3/2026
Confidence: 70%Source
benchmarks
opinion
Bullish
lab researcher

New benchmarks should measure real world impact like diseases cured, math problems solved, scientific discoveries, and materials invented rather than traditional metrics

Cristobal Valenzuela
8/3/2026
Confidence: 80%Source
benchmarks
fact
Bearish
critic

Astra did not have a control group in its evaluation

Gary Marcus
8/3/2026
Confidence: 70%Source
benchmarks
opinion
Bearish
critic

Astra is maybe not much better than Fable and certainly not ASI

Gary Marcus
8/3/2026
Confidence: 60%Source
benchmarks
opinion
Bearish
critic

Astra was a play to distract from OpenAI's increasingly terrible economics

Gary Marcus
8/3/2026
Confidence: 50%Source
benchmarks
fact
Bearish
critic

Fable and Sol can do a bunch of the same stuff as Astra, suggesting Astra is incremental rather than revolutionary

Gary Marcus
8/3/2026
Confidence: 70%Source
benchmarks
critique
Bearish
critic

OpenAI didn't include evidence on any tasks that don't involve formal verification that Astra represents significant advances over earlier models

Gary Marcus
8/3/2026
Confidence: 70%Source
benchmarks
prediction
Neutral
critic

AGI will come someday, but pretending every new model is AGI is not hastening that moment

Gary Marcus
8/3/2026
Confidence: 80%Source
benchmarks
opinion
Neutral
critic

AGI must be general and work across domains beyond those that can be formalized

Gary Marcus
8/2/2026
Confidence: 90%Source
benchmarks
critique
Bearish
critic

Astra being impressive at math alone does not qualify it as AGI

Gary Marcus
8/2/2026
Confidence: 80%Source
benchmarks
opinion
Neutral
critic

Writing good video scripts should be within the capabilities of AGI

Gary Marcus
8/2/2026
Confidence: 70%Source
benchmarks
fact
Neutral
independent

A new tool called 'smevals' enables running small eval suites against models, harnesses, and prompts

Simon Willison
8/2/2026
Confidence: 95%Source
benchmarks
opinion
Neutral
independent

Defining vocabulary for evaluation tools is one of the hardest parts of building such projects

Simon Willison
8/2/2026
Confidence: 75%Source
benchmarks
fact
Bearish
academic

13.6% of SWE-bench Verified instances exhibit PR-Issue misalignment across five patterns, undermining benchmark reliability

Artificial Intelligence
8/2/2026
Confidence: 85%Source
benchmarks
critique
Bearish
academic

SWE-bench-like benchmarks suffer from systematic misalignment due to the complexity of PR-Issue pairing in large repositories

Artificial Intelligence
8/2/2026
Confidence: 80%Source
benchmarks
fact
Neutral
academic

Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups

Machine Learning
8/2/2026
Confidence: 85%Source
14
Page 5 of 14
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,949 pending.