rlhffactbullishNeuron On-Policy Self-Distillation can enable annotation-free post-training by leveraging internal neuron activations for data selection and teacher constructionMachine Learning27 Jul 2026http://arxiv.org/abs/2607.02460v1