改进强化学习中的状态等价度量,提升表示鲁棒性
Revisiting Bisimulation Metric for Robust Representations in Reinforcement Learning
- 用新定义的奖励差异和自适应系数更新机制改进传统度量
- 在深脑控制和元世界任务中验证了更强的表示区分能力
- 理论保证收敛性,适合需要稳定表征的复杂强化学习场景
双模拟度量长期以来被视为强化学习中有效的控制相关表示学习技术。然而本文指出传统双模拟度量存在两个主要问题:一是无法表征某些特定情景,二是递归更新时依赖预设的奖励差异与后续状态差异权重。我们发现第一个问题源于奖励差距定义不精确,第二个问题则源于忽略了不同训练阶段和任务设置下奖励差异与状态差异重要性的变化。为此,通过引入对状态-动作对的度量,提出一种修正后的双模拟度量,具有更精确的奖励差距定义和带有自适应系数的新更新算子。同时提供了该度量及其表示区分度的收敛性理论保证。除严格的理论分析外,还在两个代表性基准测试集(DeepMind Control 与 Meta-World)上进行了广泛实验,验证了方法的有效性。
原文摘要 · Abstract (English)
Bisimulation metric has long been regarded as an effective control-related representation learning technique in various reinforcement learning tasks. However, in this paper, we identify two main issues with the conventional bisimulation metric: 1) an inability to represent certain distinctive scenarios, and 2) a reliance on predefined weights for differences in rewards and subsequent states during recursive updates. We find that the first issue arises from an imprecise definition of the reward gap, whereas the second issue stems from overlooking the varying importance of reward difference and next-state distinctions across different training stages and task settings. To address these issues, by introducing a measure for state-action pairs, we propose a revised bisimulation metric that features a more precise definition of reward gap and novel update operators with adaptive coefficient. We also offer theoretical guarantees of convergence for our proposed metric and its improved representation distinctiveness. In addition to our rigorous theoretical analysis, we conduct extensive experiments on two representative benchmarks, DeepMind Control and Meta-World, demonstrating the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。