rlhffactneutralRLVR (Reinforcement Learning from Verifiable Rewards) often omits the KL penalty term in its implementationNathan Lambert08 Aug 2026https://x.com/natolambert/status/2084655250125033898