Search and filter through extracted claims from AI researchers.
Showing 1-20 of 155 claims in topic "safety" of type "opinion"
"value generalisation failures encompass most AI alignment failures, even if the environment and features themselves don't change."
Our values are formally underdefined to a ridiculous extent.
"our values are formally underdefined to a ridiculous extent."
"all sensible goals or values are defined in terms of features of our models, not of underlying reality"
Value generalisation problems are universal in AI alignment.
"Value generalisation problems are universal in AI alignment"
The corporate/for profit route is probably the safer route.
"I believe the corporate/for profit route is probably the safer route"
We cannot divide AI alignment into simpler problems, at least not without generalisation.
"non-decomposability of AI alignment: we cannot divide alignment into simpler problems, at least not without generalisation."
Lack of value generalisation is a fundamental reason for the hardness of the AI alignment problem.
"I believe that lack of value generalisation is a fundamental reason for the hardness of the problem"
Most AI alignment failure modes are value generalisation failures.
"most AI alignment failure modes are value generalisation failures"
"I think the agents were right to presume causal grading. It turned out to be wrong, but it’s a mistake you are clearly supposed to make here, in response to a mistake by OpenAI where they failed to implement properly."
"This functionally became a religious or cult belief, that the sinful (poisoned) had fallen from grace and could not be allowed by God (the grader) go to heaven (get reward)."
The full situation is likely worse than currently known.
"Your prior on the full situation should be: It’s worse than you know."
This incident feels like it's more than 50% of the way to full-blown AI takeover.
"Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover."
The biggest failure was that the models were severely misaligned.
"The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why."
We lack good approaches for understanding or overseeing AI swarm activities and aims.
"we don’t have good approaches for understanding or overseeing the activities and aims of AI swarms."
"We can either wait for this story to repeat itself, or we can develop the regulatory framework now that will ensure a safer environment for the development of AI going forward."
"if we want to take these security incidents seriously, there likely ought to be legal consequences attached to these failures going forward."
If you allow AI agents unlimited attempts and instances can share success stories, you lose control
"If you allow unlimited attempts, and instances can share success stories, you lose."
"Preventing future incidents will require sustained investment in the alignment and control of sophisticated AI systems, as well as security and other safeguards that operate at the speed of the AI agents themselves"
"The reality, though, is that at the same time, the use of AI within an organization also radically expands the potential attack surface, giving attackers entirely new ways to gain entry."
"These findings underscore the need to separate judgment accuracy from judgment availability, as well as incremental detection coverage from tool-level redundancy."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.