探究最大熵强化学习在混沌系统中的抗噪鲁棒性及其理论解释
Evidence on the Regularisation Properties of Maximum-Entropy Reinforcement Learning
- 用熵正则化优化策略,提升对观测噪声的鲁棒性
- 在含高斯噪声的混沌系统中,策略泛化能力显著增强
- 通过统计学习理论复杂度指标,解释了抗噪机制
在存在高斯噪声的可观察混沌动力系统上,研究了最大熵强化学习所学策略的泛化与鲁棒性。首先观察到熵正则化策略在观测受噪声污染时仍保持稳定性;其次借用统计学习理论中的复杂度度量,解释并预测该现象。结果表明,熵正则化策略优化与抗噪声能力之间存在明确关联,该关联可通过所选复杂度度量加以描述。
原文摘要 · Abstract (English)
The generalisation and robustness properties of policies learnt through Maximum-Entropy Reinforcement Learning are investigated on chaotic dynamical systems with Gaussian noise on the observable. First, the robustness under noise contamination of the agent's observation of entropy regularised policies is observed. Second, notions of statistical learning theory, such as complexity measures on the learnt model, are borrowed to explain and predict the phenomenon. Results show the existence of a relationship between entropy-regularised policy optimisation and robustness to noise, which can be described by the chosen complexity measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。