HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claimsscaling
scaling
fact
bullish

The reward signal was the first thing to go synthetic in 2022 with InstructGPT and Constitutional AI

The first thing to go synthetic was, counterintuitively, the judge. InstructGPT established the now-canonical trick: collect human preferences once, train a reward model, and let the policy optimize against the model rather than the humans.
swyx & Alessio28 Aug 2026

https://www.latent.space/p/ainews-10-worse-100x-cheaper-10000x