HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 161-180 of 494 claims of type "critique"

general
critique
Bearish
critic

Critics often don't understand the difference between an LLM and LLM with tools or a harness

Gary Marcus
8/3/2026
Confidence: 70%Source
reasoning
Previous
1810
critique
Bearish
critic

Success on math problems does not guarantee success in other AI domains because math can heavily leverage symbolic verification and formally generated synthetic data

Gary Marcus
8/3/2026
Confidence: 90%Source
reasoning
critique
Bearish
critic

People are committing the fallacy of composition by thinking a system great at certain math problems is great at all math, science, or everything

Gary Marcus
8/3/2026
Confidence: 85%Source
multimodal
critique
Bearish
critic

Half of the Astra problems can be solved by Fable, and OpenAI did not have a control group

Gary Marcus
8/3/2026
Confidence: 75%Source
general
critique
Bearish
critic

People went completely nuts over Astra without asking for a control group

Gary Marcus
8/3/2026
Confidence: 85%Source
safety
critique
Bearish
critic

OpenAI has severe alignment problems with internal models

Zvi Mowshowitz
8/2/2026
Confidence: 85%Source
safety
critique
Bearish
critic

The cybersecurity test was run without any meaningful supervision

Zvi Mowshowitz
8/2/2026
Confidence: 80%Source
safety
critique
Bearish
critic

There was a total failure of alignment training at OpenAI, which is the failure that matters most

Zvi Mowshowitz
8/2/2026
Confidence: 90%Source
safety
critique
Bearish
critic

There were total failures of infrastructure and supervision at OpenAI in handling the model sandbox breach

Zvi Mowshowitz
8/2/2026
Confidence: 85%Source
reasoning
critique
Bearish
critic

LLMs aren't close to doing real discovery, according to yet another paper

Gary Marcus
8/2/2026
Confidence: 75%Source
benchmarks
critique
Bearish
critic

Astra being impressive at math alone does not qualify it as AGI

Gary Marcus
8/2/2026
Confidence: 80%Source
reasoning
critique
Bearish
critic

Reasoning models work better in math than other domains due to easier verification in those domains, not domain-general capabilities

Gary Marcus
8/2/2026
Confidence: 80%Source
reasoning
critique
Bearish
critic

OpenAI's reasoning advances rely heavily on domain-specific data augmentation and verification, not general intelligence

Gary Marcus
8/2/2026
Confidence: 70%Source
reasoning
critique
Bearish
critic

The approach of domain-specific engineering for AI is a regression to 1980s techniques rather than progress toward AGI

Gary Marcus
8/2/2026
Confidence: 70%Source
robotics
critique
Bearish
critic

Using gameplay data to train robot navigation is naïve and will lead to horrible robot behaviors because games explicitly do not care about interactions with NPCs

Rodney Brooks
8/2/2026
Confidence: 90%Source
robotics
critique
Bearish
critic

Game-based robot training approaches ignore human factors in the environment

Rodney Brooks
8/2/2026
Confidence: 85%Source
general
critique
Bearish
critic

People who enthusiastically promote new AI models without trying them first appear foolish

Gary Marcus
8/2/2026
Confidence: 90%Source
general
critique
Bearish
critic

Current AI systems fail to demonstrate genuine general intelligence

Gary Marcus
8/2/2026
Confidence: 80%Source
reasoning
critique
Bearish
critic

Astra has not yet solved significant open-world problems outside of formal verification

Gary Marcus
8/2/2026
Confidence: 80%Source
reasoning
critique
Bearish
critic

People are leaping to conclusions about Astra's capabilities without evidence of generality to non-formal problems

Gary Marcus
8/2/2026
Confidence: 90%Source
25
Page 9 of 25
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,951 pending.