Search and filter through extracted claims from AI researchers.
Showing 81-94 of 94 claims in topic "rlhf"
Rubrics used in RLVR are prone to over-optimization in a way similar to reward models
Llama 4 and other models are gaming leaderboards through optimization strategies
Over-optimization leads to reward hacking, sycophancy, and verbosity in language models
Over-optimization in RLHF leads to reward hacking, sycophancy, and verbosity issues
Over-optimization manifests in AI systems through reward hacking, sycophancy, and verbosity issues
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.