arXiv:2602.06939cs.LGcs.AI2026-02被引 7

用拓扑学视角解析强化学习中的时序差分信号,提升非马尔可夫环境下的学习性能。

Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics

  • 将TD误差视为状态转移空间的1-上链,揭示其拓扑结构
  • 分解出可积分与非可积部分,实现更优的近似逼近
  • 新算法HFPS在非马尔可夫场景下显著提升稳定性与表现

由于长程依赖、部分可观测性和记忆效应,现实环境中的动态常为非马尔可夫。此时贝尔曼方程仅近似成立。现有工作多聚焦于实用算法设计,对关键问题如贝尔曼框架可捕捉何种动态、如何启发最优近似算法等缺乏理论分析。本文提出一种基于拓扑学的时序差分(TD)强化学习新视角:将TD误差视为状态转移拓扑空间中的1-上链,马尔可夫动态对应拓扑可积性。通过贝尔曼-德·拉姆投影,实现TD误差的霍奇型分解——可积分成分与拓扑残差。进一步提出霍奇流策略搜索(HodgeFlow Policy Search, HFPS),拟合势能网络以最小化非可积投影残差,获得稳定性和敏感性保障。数值实验表明,HFPS在非马尔可夫环境下显著提升强化学习性能。

原文摘要 · Abstract (English)

Non-Markovian dynamics are commonly found in real-world environments due to long-range dependencies, partial observability, and memory effects. The Bellman equation that is the central pillar of Reinforcement learning (RL) becomes only approximately valid under Non-Markovian. Existing work often focus on practical algorithm designs and offer limited theoretical treatment to address key questions, such as what dynamics are indeed capturable by the Bellman framework and how to inspire new algorithm classes with optimal approximations. In this paper, we present a novel topological viewpoint on temporal-difference (TD) based RL. We show that TD errors can be viewed as 1-cochain in the topological space of state transitions, while Markov dynamics are then interpreted as topological integrability. This novel view enables us to obtain a Hodge-type decomposition of TD errors into an integrable component and a topological residual, through a Bellman-de Rham projection. We further propose HodgeFlow Policy Search (HFPS) by fitting a potential network to minimize the non-integrable projection residual in RL, achieving stability/sensitivity guarantees. In numerical evaluations, HFPS is shown to significantly improve RL performance under non-Markovian.

强化学习拓扑学非马尔可夫策略搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。