用状态熵正则提升强化学习在复杂扰动下的鲁棒性
State Entropy Regularization for Robust Reinforcement Learning
- 通过状态熵正则化增强对结构化、空间相关扰动的适应能力
- 理论证明其在奖励与转移不确定性下仍保持稳健性
- 适合关注迁移学习中环境变化鲁棒性的研究者
状态熵正则化在强化学习中表现出更优的探索效率和样本复杂度,但其理论保证尚未被深入研究。本文首次证明,该方法能有效提升对结构化及空间相关扰动的鲁棒性,这类变化在迁移学习中常见,却常被标准鲁棒强化学习方法忽略。我们给出了全面的鲁棒性分析,包括在奖励与转移不确定性下的形式化保证,以及性能下降的场景。分析对比了状态熵与广泛使用的策略熵正则化,揭示了二者不同的优势。从实践角度看,状态熵带来的鲁棒性优势对用于策略评估的回溯次数更为敏感。
原文摘要 · Abstract (English)
State entropy regularization has empirically shown better exploration and sample complexity in reinforcement learning (RL). However, its theoretical guarantees have not been studied. In this paper, we show that state entropy regularization improves robustness to structured and spatially correlated perturbations. These types of variation are common in transfer learning but often overlooked by standard robust RL methods, which typically focus on small, uncorrelated changes. We provide a comprehensive characterization of these robustness properties, including formal guarantees under reward and transition uncertainty, as well as settings where the method performs poorly. Much of our analysis contrasts state entropy with the widely used policy entropy regularization, highlighting their different benefits. Finally, from a practical standpoint, we illustrate that compared with policy entropy, the robustness advantages of state entropy are more sensitive to the number of rollouts used for policy evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。