arXiv:2607.08444stat.MLcs.LG2026-07被引 2

提出高效量化分布强化学习方法,实现最优样本效率与统计推断。

Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning

  • 基于分位数投影贝尔曼方程构造估计器,实现有限维返回分布表征。
  • 证明估计误差为 $\widetilde{O}(\sqrt{m/n})$,达到 $\sqrt{n}$ 最优收敛率。
  • 支持函数泛函的统计推断,适用于高维与无限维分位数场景。

本文从统计效率角度研究基于分位数的分布强化学习。针对分布策略评估问题,目标是刻画给定策略下回报分布(即折扣累积奖励的分布)。为获得回报分布的有限维表示,考虑由分位数投影分布贝尔曼方程诱导的分位数固定点 $η_m$。假设存在生成模型,基于经验马尔可夫决策过程构建估计器 $η_m^{(n)}$。对于固定分位数数量 $m$,在上确界 $W_\infty$ 范数下建立 $η_m^{(n)}$ 与 $η_m$ 的非渐近误差界,表明估计误差关于 $m$ 和 $n$ 的尺度为 $\widetilde{O}(\sqrt{m/n})$,说明该问题具有样本效率,可实现最优参数级 $\sqrt{n}$ 收敛速率。推导分位数参数 $\sqrt{n}(θ_m^{(n)}-θ_m)$ 的渐近分布,并刻画半参数效率界,该界被本估计器所达到。进一步研究分位数数量发散的渐近情形,刻画极限协方差结构,证明其匹配分布策略评估非参数模型的半参数效率界,表明分位数估计器在无限维极限下仍保持渐近高效。最后,为光滑泛函 $\sqrt{n}(η_m^{(n)}(s)-η_m(s))f$ 建立 Berry--Esseen 定理,为分位数投影回报分布泛函的统计有效推断提供基础。

原文摘要 · Abstract (English)

In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewards under a given policy. To obtain a finite-dimensional representation of the return distribution, we consider the quantile fixed point $η_m$ induced by the quantile-projected distributional Bellman equation. Assuming access to a generative model, we construct an estimator $η_m^{(n)}$ based on an empirical Markov decision process. For a fixed number of quantiles $m$, we establish a non-asymptotic error bound for $η_m^{(n)}$ and $η_m$ under the supremum $W_\infty$ metric, showing that the estimation error scales as $\widetilde{O}(\sqrt{m/n})$ with respect to $m$ and $n$. This implies that the quantile-based distributional policy evaluation problem can be solved with sample efficiency, achieving the optimal parametric $\sqrt{n}$ convergence rate. We derive the asymptotic distribution of the quantile parameters $\sqrt{n}(θ_m^{(n)}-θ_m)$ and characterize the semiparametric efficiency bound, which is attained by our estimator. Beyond the fixed-dimensional setting, we investigate the asymptotic regime in which the number of quantiles diverges. We characterize the limit covariance structure and show that it matches the semiparametric efficiency bound of the nonparametric model for distributional policy evaluation, showing that quantile-based estimators remain asymptotically efficient in the infinite-dimensional limit. Finally, we establish a Berry--Esseen theorem for smooth functionals $\sqrt{n}(η_m^{(n)}(s)-η_m(s))f$, thereby providing a foundation for statistically valid inference on functionals of the quantile-projected return distribution.

强化学习分布评估统计效率分位数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。