arXiv:2509.18714cs.LGcs.AI2025-09NeurIPS被引 5

提出跨MDP状态相似度度量GBSM,提升策略迁移与聚合精度

A Generalized Bisimulation Metric of State Similarity between Markov Decision Processes: From Theoretical Propositions to Applications

  • 构建跨MDP的广义双仿真度量GBSM,满足对称性与三角不等式
  • 理论证明其在策略迁移中误差上界严格优于传统双仿真度量
  • 提供闭式样本复杂度公式,适用于多MDP场景的高效估计

双仿真度量(BSM)是衡量马尔可夫决策过程(MDP)内状态相似性的有力工具,表明在BSM中距离更近的状态具有更相似的最优值函数。尽管已在强化学习中用于状态表征学习和策略探索,但在多MDP场景(如策略迁移)中的应用仍具挑战。以往工作尝试将BSM推广至多MDP,但缺乏对其数学性质的严格分析。本文首次形式化定义了跨MDP的广义双仿真度量(GBSM),并严格证明其具备三大基本性质:对称性、跨MDP三角不等式及相同状态空间下的距离有界性。基于这些性质,我们理论分析了策略迁移、状态聚合与基于采样的估计问题,获得了比标准BSM更严格的误差上界。此外,GBSM提供了闭式样本复杂度表达式,优于现有基于BSM的渐近结果。数值实验验证了理论结论,并展示了GBSM在多MDP场景中的有效性。

原文摘要 · Abstract (English)

The bisimulation metric (BSM) is a powerful tool for computing state similarities within a Markov decision process (MDP), revealing that states closer in BSM have more similar optimal value functions. While BSM has been successfully utilized in reinforcement learning (RL) for tasks like state representation learning and policy exploration, its application to multiple-MDP scenarios, such as policy transfer, remains challenging. Prior work has attempted to generalize BSM to pairs of MDPs, but a lack of rigorous analysis of its mathematical properties has limited further theoretical progress. In this work, we formally establish a generalized bisimulation metric (GBSM) between pairs of MDPs, which is rigorously proven with the three fundamental properties: GBSM symmetry, inter-MDP triangle inequality, and the distance bound on identical state spaces. Leveraging these properties, we theoretically analyse policy transfer, state aggregation, and sampling-based estimation in MDPs, obtaining explicit bounds that are strictly tighter than those derived from the standard BSM. Additionally, GBSM provides a closed-form sample complexity for estimation, improving upon existing asymptotic results based on BSM. Numerical results validate our theoretical findings and demonstrate the effectiveness of GBSM in multi-MDP scenarios.

强化学习状态相似性策略迁移度量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。