safety
critique
bearish
Most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs.
most alignment approaches fail at the moment where their fundamental assumptions about their features fail to generalise. LLMs fail similarly to other designs.
AI Alignment Forum01 Sept 2026