637 claims over the last 90 days
No lab researcher claims on this topic.
Absent generalisation, motivationally aligned AIs are not safe.
If alignment is going to work, we need to allow a certain sloppiness or imperfection on our part; generalisation grants this sloppiness.
Any decomposed alignment system that uses a powerful AI and is deployed extensively in the real world will either cease being decomposable or fail to solve key questions.
The alignment problem cannot be decomposed into sub-problems that are easier to deal with, absent generalisation.
Most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs.
Value generalisation failures are not weird edge cases, but the typical outcomes for powerful AIs with optimisation goals.
Value generalisation failures encompass most AI alignment failures, even if the environment and features themselves don't change.
Our values are formally underdefined to a ridiculous extent.
Value generalisation will be crucial for AI alignment.
Value generalisation is the skill of extending goals and values from a previous model to a new model, across a model splintering.
No other claims on this topic.