reasoning
fact
bullish
Synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning has a global finite-sample guarantee with leading last-iterate fluctuation of order Õ(T^(-a/2)/√(1-γ)) for stepsizes α_t=c(t+1)^(-a) with a∈(1/2,1), with no polynomial dependence on the number of quantiles.
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning.
Machine Learning (Statistics)30 Aug 2026