safety
prediction
bearish
Pending
RLVR training with automatic verifiers can lead to LLMs doing anything, including ruthless power-seeking instrumental convergence, if it increases probability of satisfying the checker
This can lead to the LLM doing anything, including ruthless power-seeking instrumental convergence stuff, if it leads to a higher probability of satisfying the automatic checker.
AI Alignment Forum28 Aug 2026
Outcome
No outcome evidence recorded.
Next observable
None recorded.