rlhfopinionbullishFine-tuning with expert judgment data can beat prompting-only approaches by a significant margin, even as general-purpose models improveJohn Schulman27 Jul 2026https://x.com/johnschulman2/status/2072132258463719658