arXiv:2505.19058cs.LGmath.OC2025-05被引 3

针对状态转移不确定的强化学习,提出抗扰动的Q学习新算法。

Distributionally Robust Deep Q-Learning

  • 用Sinkhorn距离正则化贝尔曼算子,处理模型不确定性。
  • 在S&P 500投资组合优化中,比传统方法更稳定且收益更高。
  • 适合高风险场景下的鲁棒决策,如金融、机器人控制。

我们提出一种新型分布鲁棒深度Q学习算法,适用于连续状态空间的非表格型问题,其中马尔可夫决策过程的状态转移存在模型不确定性。通过考虑参考概率测度邻域内最坏情况的转移来建模不确定性。为求解最坏情况下的最优策略,将非线性贝尔曼方程进行对偶化和正则化,使用Sinkhorn距离实现,并由深度神经网络参数化。该方法可改造Deep Q-Network以优化最坏情况下的状态转移。通过多个应用验证了其可计算性与有效性,包括基于S&P 500数据的投资组合优化任务。

原文摘要 · Abstract (English)

We propose a novel distributionally robust $Q$-learning algorithm for the non-tabular case accounting for continuous state spaces where the state transition of the underlying Markov decision process is subject to model uncertainty. The uncertainty is taken into account by considering the worst-case transition from a ball around a reference probability measure. To determine the optimal policy under the worst-case state transition, we solve the associated non-linear Bellman equation by dualising and regularising the Bellman operator with the Sinkhorn distance, which is then parameterized with deep neural networks. This approach allows us to modify the Deep Q-Network algorithm to optimise for the worst case state transition. We illustrate the tractability and effectiveness of our approach through several applications, including a portfolio optimisation task based on S\&{P}~500 data.

强化学习鲁棒优化深度Q网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。