reasoningfactneutralCore ideas for reasoning improvement through iterative training with rejection sampling were already present before 2022, but what changed was using natural language instead of programming languagesDenny Zhou27 Jul 2026https://x.com/denny_zhou/status/2081521217807520004