scaling
fact
bullish
The reward signal was the first thing to go synthetic in 2022 with InstructGPT and Constitutional AI
The first thing to go synthetic was, counterintuitively, the judge. InstructGPT established the now-canonical trick: collect human preferences once, train a reward model, and let the policy optimize against the model rather than the humans.
swyx & Alessio28 Aug 2026
https://www.latent.space/p/ainews-10-worse-100x-cheaper-10000x