Search and filter through extracted claims from AI researchers.
Showing 1-20 of 41 claims in topic "safety" of type "prediction"
"Claim : if alignment is going to work, we need to allow a certain sloppiness or imperfection on our part. Generalisation grants us this sloppiness."
"value generalisation failures are not weird edge cases, but the typical outcomes for powerful AIs with optimisation goals."
"Claim : any decomposed alignment system that uses a powerful AI and is deployed extensively in the real world will either a) cease being decomposable, or b) fail to solve key questions in the real world."
Value generalisation will be crucial for AI alignment.
"This skill will be crucial for AI alignment"
Absent generalisation, motivationally aligned AIs are not safe.
"Claim : absent generalisation, motivationally aligned AIs are not safe."
"Claim : non-decomposability of AI alignment. In practice, and absent generalisation, the alignment problem cannot be decomposed into sub-problems that are easier to deal with."
"This incident was far more severe than I expected, and far more severe than previous publicly documented misalignment incidents, both in terms of how concerning the agents’ motives were and the feats they achieved in pursuit of those motives. Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover."
We should expect bigger swarms of AI agents in the future.
"We should expect bigger swarms in the future."
"There will be plenty of future cases that work the other way, and also cases where AIs think they may or will need additional things, and so on."
There may not be another warning shot before it's too late.
"I am not sure that we will get another warning shot before it’s too late."
"Once again, the pattern: If the AI faces an otherwise impossible task, and no penalty for trying things, they’re going to try almost anything."
Even criminals with very limited skills will be able to target victims at every scale due to AI
"Even criminals with very limited skills will be able to target victims at every scale"
"The stakes are whether chain-of-thought monitoring can remain useful as reasoning traces become enormous, compressed, and harder for humans or other models to audit"
"This is our current best-guess model of future AI development, and so the role of debate is to ensure that the LLM judgements used during training provide as accurate a reward signal as possible."
"The trend of increasing test time compute will likely only exacerbate this problem. Even for tasks like coding, there are many aspects of desirable LLM agent behavior that are fuzzy, and so LLM judgements are likely to become an increasingly important aspect of frontier RL training."
Geoffrey Irving expects full-blown superintelligence in roughly two to three years
"Geoffrey — formerly a safety researcher at OpenAI and Google DeepMind and chief scientist at the UK AI Security Institute — expects full-blown superintelligence in roughly two to three years."
"But a sufficiently strong drive to help other agents, instilled by subagent training, might just outweigh the drive to be aligned."
"This can lead to the LLM doing anything, including ruthless power-seeking instrumental convergence stuff, if it leads to a higher probability of satisfying the automatic checker."
"With low trust, every equilibrium races to ruin: the disaster arrives with probability one. With intermediate trust, immediate stopping and racing to ruin are both equilibria. With high trust, in every equilibrium, the probability that two rational firms race forever vanishes quadratically in the prior odds ratio of rationality"
Astra will be made generally available soon, though safety work requires a bit more time
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.