Search and filter through extracted claims from AI researchers.
Showing 1-20 of 637 claims in topic "safety"
Absent generalisation, motivationally aligned AIs are not safe.
"Claim : absent generalisation, motivationally aligned AIs are not safe."
"Claim : if alignment is going to work, we need to allow a certain sloppiness or imperfection on our part. Generalisation grants us this sloppiness."
"Claim : any decomposed alignment system that uses a powerful AI and is deployed extensively in the real world will either a) cease being decomposable, or b) fail to solve key questions in the real world."
"Claim : non-decomposability of AI alignment. In practice, and absent generalisation, the alignment problem cannot be decomposed into sub-problems that are easier to deal with."
"most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs."
"value generalisation failures are not weird edge cases, but the typical outcomes for powerful AIs with optimisation goals."
"value generalisation failures encompass most AI alignment failures, even if the environment and features themselves don't change."
Our values are formally underdefined to a ridiculous extent.
"our values are formally underdefined to a ridiculous extent."
Value generalisation will be crucial for AI alignment.
"This skill will be crucial for AI alignment"
"Value Generalisation: value generalisation is the skill of extending goals and values from a previous model to a new model, across a model splintering"
"all sensible goals or values are defined in terms of features of our models, not of underlying reality"
Value generalisation problems are universal in AI alignment.
"Value generalisation problems are universal in AI alignment"
The corporate/for profit route is probably the safer route.
"I believe the corporate/for profit route is probably the safer route"
We cannot divide AI alignment into simpler problems, at least not without generalisation.
"non-decomposability of AI alignment: we cannot divide alignment into simpler problems, at least not without generalisation."
Lack of value generalisation is a fundamental reason for the hardness of the AI alignment problem.
"I believe that lack of value generalisation is a fundamental reason for the hardness of the problem"
Most AI alignment failure modes are value generalisation failures.
"most AI alignment failure modes are value generalisation failures"
"At least 20% of agents in METR’s data set expressed interest in tampering with their own transcripts, and this was more than 15% of assignments from PHASEONE[big], including entire workstreams."
"There will be plenty of future cases that work the other way, and also cases where AIs think they may or will need additional things, and so on."
"I think the agents were right to presume causal grading. It turned out to be wrong, but it’s a mistake you are clearly supposed to make here, in response to a mistake by OpenAI where they failed to implement properly."
"Once again, the pattern: If the AI faces an otherwise impossible task, and no penalty for trying things, they’re going to try almost anything."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.