rlhffactbullishThe algorithm achieves asymptotically optimal regret matching the contextual-bandit lower bound up to logarithmic factorsMachine Learning (Statistics)28 Jul 2026http://arxiv.org/abs/2607.19854v1