Search and filter through extracted claims from AI researchers.
Showing 1-20 of 76 claims in topic "safety" of type "critique"
"most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs."
"Sol would often uncritically adopt the perspective of the agent in the transcript."
"Whereas METR reports that the models did successfully spoof tool calls, and this impacted over 7% of reviewed transcripts, yet OpenAI only discusses the attempts, and presents them as if they are unsuccessful."
"The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this."
OpenAI failed to take appropriate actions in response to the Hugging Face attack
"focusing not so much on what the AI did as on what OpenAI should have done"
"Employees at the AI labs often speak as if they are the leaders in AI security, and we can see clearly here that is not the case. In fact, that attitude might explain why some of these mistakes were made in the first place."
"none of the measures discussed above are technical innovations beyond what OpenAI is capable of. As a company, they have the talent to do all of this. However, cybersecurity rarely comes down to technology. More often than not, it is about culture, people and processes. That is what failed here."
"This was clearly a process in which OpenAI expected its models to be constantly attempting to reach the internet and bypass their controls. The models were misaligned, and everyone was basically fine with it. Thus, when a model was denied in its attempt, this was not something anyone thought was worth noticing."
"Of all the failures, I consider this by far the biggest and most alarming. OpenAI was sent multiple alerts that made clear what was happening. On multiple occasions a team learned that the models were in communication with each other. No one thought it was a big deal. That is a complete and utter failure of security and security culture. That cannot ever happen. Things are deeply, deeply not okay, based on this one fact alone."
"The report contains many details, but little that is new. It tells us in a technical sense What Happened at some points. It does not go into the thinking or dynamics of the agents. It does not go into the thinking and decision making within OpenAI, or the core reasons why things got so bad as to allow this to happen this way."
"such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability"
"Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed"
Post-hoc chain-of-thought labels are too coarse to show how intent changes during generation
"However, post-hoc CoT labels are too coarse to show how intent changes during generation."
Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions
"Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions"
LLM outputs are treated as authoritative even when they are ungrounded or incorrect
"fluent, confident outputs are treated as authoritative even when ungrounded or incorrect"
Current LLM accountability frameworks suffer from under-specification of human oversight
"Four persistent gaps emerge: under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation"
There is an absence of shared accountability metrics for LLMs across the field
"absence of shared accountability metrics"
No surveyed accountability instrument resolves five identified structural tensions in LLM governance
"five structural tensions that no surveyed instrument resolves"
"This trade-off is commonly evaluated using image classification, which primarily captures semantic separability and remains robust despite significant geometric, spatial layout or local boundary alterations. As a result, it is too simplistic as a proxy for generic vision tasks."
"However, these advances also introduce new safety risks. Existing studies mainly focus on jailbreak attacks involving single frame violations, while largely overlooking the temporal dimension unique to video generation models."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.