rlhffactbullishA new algorithm achieves horizon-free regret of O(sqrt(SAK)+S^8A^3) for finite-horizon tabular MDPs, completely removing log H dependenceMachine Learning (Statistics)28 Jul 2026http://arxiv.org/abs/2607.19854v1