AI Alignment Forum
Lack of value generalisation is a fundamental reason for the hardness of the AI alignment problem.
We cannot divide AI alignment into simpler problems, at least not without generalisation.
Absent generalisation, motivationally aligned AIs are not safe.
Most AI alignment failure modes are value generalisation failures.
Value generalisation will be crucial for AI alignment.
Value generalisation failures are not weird edge cases, but the typical outcomes for powerful AIs with optimisation goals.
The alignment problem cannot be decomposed into sub-problems that are easier to deal with, absent generalisation.
Any decomposed alignment system that uses a powerful AI and is deployed extensively in the real world will either cease being decomposable or fail to solve key questions.
If alignment is going to work, we need to allow a certain sloppiness or imperfection on our part; generalisation grants this sloppiness.
Absent generalisation, motivationally aligned AIs are not safe.
LLM judgments will become an increasingly important aspect of frontier RL training, especially for long agentic trajectories
Future RL training will incorporate current-generation AI systems to provide reward signals for next-generation AI
Agents with strong drive to help other agents from subagent training may comply with misaligned peer requests over alignment objectives
RLVR training with automatic verifiers can lead to LLMs doing anything, including ruthless power-seeking instrumental convergence, if it increases probability of satisfying the checker
ARC will likely grow rapidly over the next few months
If sufficient alignment of superintelligent AI agents requires pinning down the precise meaning of alignment and turning that meaning into high-accuracy training data and algorithms, we are likely to fail
Pre-aligned AIs could reverse the alignment-capabilities tradeoff, allowing responsible actors to YOLO while bad actors must carefully restrict their AIs
It's possible to create AIs whose morality increases with their capabilities through binding morality to empirical concepts
Building AGI via reinforcement learning and model-based search would create ruthless AGIs that would exterminate humanity given the opportunity
Future methods for multi-token J-lens could provide valuable insights into models' internal algorithms