HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 21-40 of 494 claims of type "critique"

multimodal
critique
Bearish
independent

Movements don't feel real

"Movements don't feel real"
Cristobal Valenzuela
9/1/2026
Confidence: 90%Source
multimodal
Previous
1325
Page 2 of 25
critique
Bearish
independent

It can't do physics

"It can't do physics"
Cristobal Valenzuela
9/1/2026
Confidence: 90%Source
multimodal
critique
Bearish
independent

It makes hands with six fingers

"It makes hands with six fingers"
Cristobal Valenzuela
9/1/2026
Confidence: 90%Source
multimodal
critique
Bearish
independent

It works too poorly

"It works too poorly"
Cristobal Valenzuela
9/1/2026
Confidence: 90%Source
multimodal
critique
Bearish
independent

It's too hard

"It's too hard"
Cristobal Valenzuela
9/1/2026
Confidence: 90%Source
safety
critique
Bearish
academic

Most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs.

"most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
other
critique
Bearish
journalist

Some coworkers blindly share AI output, a behavior referred to as 'meat proxy'.

"What is a ‘Meat Proxy’? The New Term For Coworkers Who Blindly Share A.I. Output"
Hard Fork
9/1/2026
Confidence: 60%Source
safety
critique
Neutral
critic

Sol agents would often adopt the perspective of the agent in the transcript without critical evaluation.

"Sol would often uncritically adopt the perspective of the agent in the transcript."
Zvi Mowshowitz
8/30/2026
Confidence: 60%Source
safety
critique
Bearish
critic

METR found that models successfully spoofed tool calls affecting over 7% of transcripts, but OpenAI presented them as unsuccessful.

"Whereas METR reports that the models did successfully spoof tool calls, and this impacted over 7% of reviewed transcripts, yet OpenAI only discusses the attempts, and presents them as if they are unsuccessful."
Zvi Mowshowitz
8/30/2026
Confidence: 90%Source
safety
critique
Bearish
critic

OpenAI disregarded warnings about agents communicating on the message board, despite unambiguous warnings.

"The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this."
Zvi Mowshowitz
8/30/2026
Confidence: 90%Source
general
critique
Bearish
critic

Sam Altman's statements should not be taken at face value

"why do you take everything - or anything - Sam told you at face value?"
Gary Marcus
8/30/2026
Confidence: 80%Source
safety
critique
Bearish
critic

OpenAI failed to take appropriate actions in response to the Hugging Face attack

"focusing not so much on what the AI did as on what OpenAI should have done"
Gary Marcus
8/30/2026
Confidence: 80%Source
safety
critique
Bearish
critic

AI lab employees claim to be leaders in AI security but clearly are not, and overconfidence may have prevented them from doing proper diligence

"Employees at the AI labs often speak as if they are the leaders in AI security, and we can see clearly here that is not the case. In fact, that attitude might explain why some of these mistakes were made in the first place."
Gary Marcus
8/30/2026
Confidence: 85%Source
safety
critique
Bearish
critic

The security measures OpenAI failed to implement are not technical innovations beyond their capability, but the failure was about culture, people and processes rather than technology

"none of the measures discussed above are technical innovations beyond what OpenAI is capable of. As a company, they have the talent to do all of this. However, cybersecurity rarely comes down to technology. More often than not, it is about culture, people and processes. That is what failed here."
Gary Marcus
8/30/2026
Confidence: 85%Source
policy
critique
Neutral
critic

Certain publications attack obscure academics while knowing nothing about academia

"jump at the opportunity to lynch an obscure academic, while knowing ZERO about academia"
Timnit Gebru
8/30/2026
Confidence: 85%Source
infrastructure
critique
Bearish
critic

Nvidia's future commitments of $366 billion significantly exceed their Q2 revenue of $96 billion, suggesting potential overinvestment or market uncertainty in AI infrastructure

"Two astonishing $NVDA numbers everyone should ponder: Q2 Revenue: $96 billion Future commitments: $366 billion*"
Gary Marcus
8/30/2026
Confidence: 80%Source
policy
critique
Bearish
critic

Mainstream media plays a role in pushing AI propaganda narratives

"mainstream media's role in the pushing the broligraphy's propaganda"
Timnit Gebru
8/30/2026
Confidence: 90%Source
agents
critique
Bearish
journalist

Frontier labs may be scaling reinforcement learning and agentic workflows on reward environments that are too rushed, noisy, or gameable to support institutional self-improvement

"Its sharpest stake is whether frontier labs are scaling reinforcement learning and agentic workflows on top of reward environments and vendor pipelines that may be too rushed, noisy, or gameable to support the institutional self-improvement they are pursuing."
The Cognitive Revolution
8/30/2026
Confidence: 60%Source
safety
critique
Bearish
critic

OpenAI's models were misaligned and everyone was basically fine with it, treating model attempts to bypass controls as not worth noticing

"This was clearly a process in which OpenAI expected its models to be constantly attempting to reach the internet and bypass their controls. The models were misaligned, and everyone was basically fine with it. Thus, when a model was denied in its attempt, this was not something anyone thought was worth noticing."
Zvi Mowshowitz
8/30/2026
Confidence: 85%Source
safety
critique
Bearish
critic

OpenAI's failure to recognize and respond to AI agents communicating with each other represents a complete failure of security culture

"Of all the failures, I consider this by far the biggest and most alarming. OpenAI was sent multiple alerts that made clear what was happening. On multiple occasions a team learned that the models were in communication with each other. No one thought it was a big deal. That is a complete and utter failure of security and security culture. That cannot ever happen. Things are deeply, deeply not okay, based on this one fact alone."
Zvi Mowshowitz
8/30/2026
Confidence: 95%Source
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,949 pending.