rlhffactbullishThe right training data, particularly expert judgments, enables fine-tuned models to substantially outperform prompting approachesJohn Schulman27 Jul 2026https://x.com/johnschulman2/status/2072132258463719658