safety
prediction
neutral
Pending
Future RL training will incorporate current-generation AI systems to provide reward signals for next-generation AI
This is our current best-guess model of future AI development, and so the role of debate is to ensure that the LLM judgements used during training provide as accurate a reward signal as possible.
AI Alignment Forum28 Aug 2026
Outcome
No outcome evidence recorded.
Next observable
None recorded.