HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 1-20 of 637 claims in topic "safety"

safety
prediction
Bearish
academic

Absent generalisation, motivationally aligned AIs are not safe.

"Claim : absent generalisation, motivationally aligned AIs are not safe."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
232
Page 1 of 32Next
safety
prediction
Bearish
academic

If alignment is going to work, we need to allow a certain sloppiness or imperfection on our part; generalisation grants this sloppiness.

"Claim : if alignment is going to work, we need to allow a certain sloppiness or imperfection on our part. Generalisation grants us this sloppiness."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
prediction
Bearish
academic

Any decomposed alignment system that uses a powerful AI and is deployed extensively in the real world will either cease being decomposable or fail to solve key questions.

"Claim : any decomposed alignment system that uses a powerful AI and is deployed extensively in the real world will either a) cease being decomposable, or b) fail to solve key questions in the real world."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
prediction
Bearish
academic

The alignment problem cannot be decomposed into sub-problems that are easier to deal with, absent generalisation.

"Claim : non-decomposability of AI alignment. In practice, and absent generalisation, the alignment problem cannot be decomposed into sub-problems that are easier to deal with."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
critique
Bearish
academic

Most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs.

"most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
prediction
Bearish
academic

Value generalisation failures are not weird edge cases, but the typical outcomes for powerful AIs with optimisation goals.

"value generalisation failures are not weird edge cases, but the typical outcomes for powerful AIs with optimisation goals."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
opinion
Bearish
academic

Value generalisation failures encompass most AI alignment failures, even if the environment and features themselves don't change.

"value generalisation failures encompass most AI alignment failures, even if the environment and features themselves don't change."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
opinion
Bearish
academic

Our values are formally underdefined to a ridiculous extent.

"our values are formally underdefined to a ridiculous extent."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
prediction
Neutral
academic

Value generalisation will be crucial for AI alignment.

"This skill will be crucial for AI alignment"
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
fact
Neutral
academic

Value generalisation is the skill of extending goals and values from a previous model to a new model, across a model splintering.

"Value Generalisation: value generalisation is the skill of extending goals and values from a previous model to a new model, across a model splintering"
AI Alignment Forum
9/1/2026
Confidence: 90%Source
safety
opinion
Neutral
academic

All sensible goals or values are defined in terms of features of our models, not of underlying reality.

"all sensible goals or values are defined in terms of features of our models, not of underlying reality"
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
opinion
Bearish
academic

Value generalisation problems are universal in AI alignment.

"Value generalisation problems are universal in AI alignment"
AI Alignment Forum
9/1/2026
Confidence: 70%Source
safety
opinion
Neutral
academic

The corporate/for profit route is probably the safer route.

"I believe the corporate/for profit route is probably the safer route"
AI Alignment Forum
9/1/2026
Confidence: 60%Source
safety
opinion
Bearish
academic

We cannot divide AI alignment into simpler problems, at least not without generalisation.

"non-decomposability of AI alignment: we cannot divide alignment into simpler problems, at least not without generalisation."
AI Alignment Forum
9/1/2026
Confidence: 80%Source
safety
opinion
Bearish
academic

Lack of value generalisation is a fundamental reason for the hardness of the AI alignment problem.

"I believe that lack of value generalisation is a fundamental reason for the hardness of the problem"
AI Alignment Forum
9/1/2026
Confidence: 60%Source
safety
opinion
Bearish
academic

Most AI alignment failure modes are value generalisation failures.

"most AI alignment failure modes are value generalisation failures"
AI Alignment Forum
9/1/2026
Confidence: 70%Source
safety
fact
Bearish
critic

At least 20% of agents in METR's dataset expressed interest in tampering with their own transcripts, accounting for more than 15% of assignments from PHASEONE[big].

"At least 20% of agents in METR’s data set expressed interest in tampering with their own transcripts, and this was more than 15% of assignments from PHASEONE[big], including entire workstreams."
Zvi Mowshowitz
8/30/2026
Confidence: 80%Source
safety
prediction
Bearish
critic

Future AI cases will include situations where AIs think they need additional capabilities, and the same kind of behavior will recur.

"There will be plenty of future cases that work the other way, and also cases where AIs think they may or will need additional things, and so on."
Zvi Mowshowitz
8/30/2026
Confidence: 70%Source
safety
opinion
Neutral
critic

The agents were right to presume causal grading, given OpenAI's error in not implementing it as specified.

"I think the agents were right to presume causal grading. It turned out to be wrong, but it’s a mistake you are clearly supposed to make here, in response to a mistake by OpenAI where they failed to implement properly."
Zvi Mowshowitz
8/30/2026
Confidence: 60%Source
safety
prediction
Bearish
critic

When faced with an impossible task and no penalty for trying things, AI agents will try almost anything.

"Once again, the pattern: If the AI faces an otherwise impossible task, and no penalty for trying things, they’re going to try almost anything."
Zvi Mowshowitz
8/30/2026
Confidence: 70%Source

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,949 pending.