multimodalcritiquebearishTrajectory-level rewards in on-policy distillation cannot determine whether a failed answer arose from perception or subsequent reasoningArtificial Intelligence02 Aug 2026http://arxiv.org/abs/2607.28336v1