rlhf
prediction
neutral
Pending
Problems from controlling reward model overoptimization will rhyme with future problems of controlling rubrics for agents
Nathan Lambert29 Jul 2026
https://x.com/natolambert/status/2082188013162144066
Outcome
No outcome evidence recorded.
Next observable
None recorded.