rlhfcritiquebearishExtending order-optimal convergence guarantees to neural critics in average-reward CMDPs has remained an open problem due to a fundamental bias-cost trade-offMachine Learning02 Aug 2026http://arxiv.org/abs/2607.28390v1