rlhf
prediction
bearish
Pending
Rubrics will be prone to over-optimization similar to reward models, with RLVR being its own distinct phenomenon
Nathan Lambert26 Jul 2026
https://x.com/natolambert/status/2081018219746468088
Outcome
No outcome evidence recorded.
Next observable
None recorded.