arXiv:2605.08182cs.LGcs.AI2026-05

解决分布强化学习中量化估计的退化问题,提升风险敏感任务表现。

Quantile Geometry Regularization for Distributional Reinforcement Learning

论文配图:Quantile Geometry Regularization for Distributional Reinforcement Learning
图 1 · 摘自论文原文
  • 从采样分位数估计角度重构IQN损失,引入鲁棒性修正。
  • 在Atari游戏和风险导航任务中超越现有量化分布强化学习方法。
  • 无需修改价值目标,可直接防止分布退化,适合高风险决策场景。

基于分位数的分布强化学习通过采样分位数回归学习回报分布,但其自举目标分位数可能引发分布估计失真或退化。本文提出鲁棒分位数隐式分位数网络(RQIQN),一种轻量级的Wasserstein分布鲁棒增强方法,从分位数估计视角出发。我们首先将IQN损失的一个快照重新解释为对采样当前分位数的局部经验分位数估计问题集合。随后,通过引入Wasserstein分布鲁棒分位数估计公式,对每个局部槽进行鲁棒化处理,得到闭式、分位数依赖的目标修正项。该修正直接缓解分布退化:其中位数反对称性保持风险中性分位数平均,单调性扩大上下分位数间距,对抗分布塌陷。RQIQN在不改变底层价值目标且无需额外样本重建的情况下,正则化分位数几何结构。实验表明,RQIQN在风险敏感导航与Atari游戏任务中优于其他现有量化分布强化学习算法。

原文摘要 · Abstract (English)

Quantile-based distributional reinforcement learning methods learn return distributions through sampled quantile regression, but their bootstrapped target quantiles may induce distorted or degenerate distribution estimates. We propose Robust Quantile-based Implicit Quantile Networks (RQIQN), a lightweight Wasserstein distributionally robust enhancement boosted from a quantile estimation perspective. We first reinterpret a snapshot of IQN loss as a collection of local empirical quantile estimation problems over sampled current fractions. We then robustify each local slot with a Wasserstein distributionally robust quantile estimation formulation, yielding a closed-form, fraction-dependent correction to the Bellman target. This correction directly addresses distributional degeneration: its median antisymmetry preserves the risk-neutral quantile average, while its monotonicity enlarges upper-lower quantile gaps and counteracts collapsed distributional spread. RQIQN thus regularizes quantile geometry without changing the underlying value objective or requiring additional sample set reconstruction. Finally, we empirically show that the proposed RQIQN outperforms other existing quantile-based distributional reinforcement learning algorithms in risk-sensitive navigation and Atari games.

分布强化学习分位数回归鲁棒优化风险敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。