证明熵正则化能提升连续时间强化学习的鲁棒性
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning

- 通过理论分析建立熵正则化的鲁棒性保障
- 正则强度越大,鲁棒集扩张越明显,实验验证优于基线
- 适用于对环境扰动敏感的连续控制任务
熵正则化在连续时间强化学习中广泛用于降低对环境扰动的敏感性,但其鲁棒性优势缺乏严格的理论支撑。本文首次为熵正则化的连续时间马尔可夫决策过程建立了鲁棒性保证。我们证明,最大化熵正则目标等价于对联合奖励与转移扰动下的最坏情况强化学习问题提供下界。理论上刻画了诱导出的鲁棒集,并证明其随正则化强度单调扩张,解释了强熵正则化提升鲁棒性的经验现象。与以往离散时间分析不同,本方法消除了难以处理的状态分布熵项,且保证不依赖动作频率。在排队网络控制和做市任务上的实验验证了理论,显示熵正则策略在动态扰动下优于贪婪和ε-贪婪基线。
原文摘要 · Abstract (English)
Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigorous theoretical foundation. This paper establishes the first robustness guarantees for entropy-regularized continuous-time Markov decision processes. We show that maximizing an entropy-regularized objective yields a lower bound on a worst-case robust RL problem with joint reward and transition perturbations. We analytically characterize the induced robust sets and prove that they expand monotonically with the regularization strength, justifying the empirical observation that stronger entropy improves robustness. In contrast to prior discrete-time analyses, our results remove the intractable state-distribution entropy term and provide guarantees invariant to action frequency. Experiments on queueing network control and market making confirm our theory, showing that entropy-regularized policies outperform greedy and $ε$-greedy baselines under dynamics perturbations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。