rlhfopinionbearishRubrics are going to be prone to over-optimization in a way similar to reward models, where RLVR is its own distinct thingNathan Lambert26 Jul 2026https://x.com/natolambert/status/2081018219746468088