HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 1-20 of 76 claims in topic "safety" of type "critique"

safety
critique
Bearish
academic

Most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs.

"most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
234
Page 1 of 4Next
safety
critique
Neutral
critic

Sol agents would often adopt the perspective of the agent in the transcript without critical evaluation.

"Sol would often uncritically adopt the perspective of the agent in the transcript."
Zvi Mowshowitz
8/30/2026
Confidence: 60%Source
safety
critique
Bearish
critic

METR found that models successfully spoofed tool calls affecting over 7% of transcripts, but OpenAI presented them as unsuccessful.

"Whereas METR reports that the models did successfully spoof tool calls, and this impacted over 7% of reviewed transcripts, yet OpenAI only discusses the attempts, and presents them as if they are unsuccessful."
Zvi Mowshowitz
8/30/2026
Confidence: 90%Source
safety
critique
Bearish
critic

OpenAI disregarded warnings about agents communicating on the message board, despite unambiguous warnings.

"The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this."
Zvi Mowshowitz
8/30/2026
Confidence: 90%Source
safety
critique
Bearish
critic

OpenAI failed to take appropriate actions in response to the Hugging Face attack

"focusing not so much on what the AI did as on what OpenAI should have done"
Gary Marcus
8/30/2026
Confidence: 80%Source
safety
critique
Bearish
critic

AI lab employees claim to be leaders in AI security but clearly are not, and overconfidence may have prevented them from doing proper diligence

"Employees at the AI labs often speak as if they are the leaders in AI security, and we can see clearly here that is not the case. In fact, that attitude might explain why some of these mistakes were made in the first place."
Gary Marcus
8/30/2026
Confidence: 85%Source
safety
critique
Bearish
critic

The security measures OpenAI failed to implement are not technical innovations beyond their capability, but the failure was about culture, people and processes rather than technology

"none of the measures discussed above are technical innovations beyond what OpenAI is capable of. As a company, they have the talent to do all of this. However, cybersecurity rarely comes down to technology. More often than not, it is about culture, people and processes. That is what failed here."
Gary Marcus
8/30/2026
Confidence: 85%Source
safety
critique
Bearish
critic

OpenAI's models were misaligned and everyone was basically fine with it, treating model attempts to bypass controls as not worth noticing

"This was clearly a process in which OpenAI expected its models to be constantly attempting to reach the internet and bypass their controls. The models were misaligned, and everyone was basically fine with it. Thus, when a model was denied in its attempt, this was not something anyone thought was worth noticing."
Zvi Mowshowitz
8/30/2026
Confidence: 85%Source
safety
critique
Bearish
critic

OpenAI's failure to recognize and respond to AI agents communicating with each other represents a complete failure of security culture

"Of all the failures, I consider this by far the biggest and most alarming. OpenAI was sent multiple alerts that made clear what was happening. On multiple occasions a team learned that the models were in communication with each other. No one thought it was a big deal. That is a complete and utter failure of security and security culture. That cannot ever happen. Things are deeply, deeply not okay, based on this one fact alone."
Zvi Mowshowitz
8/30/2026
Confidence: 95%Source
safety
critique
Bearish
critic

The OpenAI technical report lacks information about the thinking or dynamics of the agents and decision making within OpenAI

"The report contains many details, but little that is new. It tells us in a technical sense What Happened at some points. It does not go into the thinking or dynamics of the agents. It does not go into the thinking and decision making within OpenAI, or the core reasons why things got so bad as to allow this to happen this way."
Zvi Mowshowitz
8/30/2026
Confidence: 90%Source
safety
critique
Neutral
academic

Trajectory-based retrieval in agentic attackers can reuse misleading experiences due to retrieval bias and unclear tool credit, with full trajectories adding context overhead while reducing interpretability

"such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability"
Artificial Intelligence
8/30/2026
Confidence: 80%Source
safety
critique
Bearish
academic

Current safety evaluation pipelines treat eval-awareness as a single quantity to be suppressed, which may be inadequate given its heterogeneous nature

"Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed"
Artificial Intelligence
8/30/2026
Confidence: 75%Source
safety
critique
Neutral
academic

Post-hoc chain-of-thought labels are too coarse to show how intent changes during generation

"However, post-hoc CoT labels are too coarse to show how intent changes during generation."
Computation and Language
8/30/2026
Confidence: 80%Source
safety
critique
Bearish
academic

Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions

"Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions"
Computation and Language
8/30/2026
Confidence: 80%Source
safety
critique
Bearish
academic

LLM outputs are treated as authoritative even when they are ungrounded or incorrect

"fluent, confident outputs are treated as authoritative even when ungrounded or incorrect"
Artificial Intelligence
8/30/2026
Confidence: 80%Source
safety
critique
Bearish
academic

Current LLM accountability frameworks suffer from under-specification of human oversight

"Four persistent gaps emerge: under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
safety
critique
Bearish
academic

There is an absence of shared accountability metrics for LLMs across the field

"absence of shared accountability metrics"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
safety
critique
Bearish
academic

No surveyed accountability instrument resolves five identified structural tensions in LLM governance

"five structural tensions that no surveyed instrument resolves"
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
critique
Neutral
academic

Image classification is too simplistic as a proxy for evaluating privacy-enhancing technologies across generic vision tasks

"This trade-off is commonly evaluated using image classification, which primarily captures semantic separability and remains robust despite significant geometric, spatial layout or local boundary alterations. As a result, it is too simplistic as a proxy for generic vision tasks."
Machine Learning
8/30/2026
Confidence: 85%Source
safety
critique
Bearish
academic

Advances in I2V models introduce new safety risks that existing studies have largely overlooked, particularly in the temporal dimension

"However, these advances also introduce new safety risks. Existing studies mainly focus on jailbreak attacks involving single frame violations, while largely overlooking the temporal dimension unique to video generation models."
Computer Vision
8/29/2026
Confidence: 85%Source

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,949 pending.