让智能体提前预判路径,提升复杂环境下的决策稳定性。
Anticipatory Reinforcement Learning: From Generative Path-Laws to Distributional Value Functions
- 用路径签名扩展状态空间,捕捉历史依赖性。
- 单轨迹下实现未来路径的确定性评估,降低计算开销。
- 适用于高波动连续时间环境,适合风险敏感型决策场景。
本文提出前瞻强化学习(ARL),旨在弥合非马尔可夫决策过程与经典强化学习架构之间的差距,尤其在仅有一个观测轨迹的约束下。在跳跃扩散和结构突变环境中,传统基于状态的方法难以捕捉准确预测所需的路径依赖几何结构。为此,我们将状态空间提升至签名增强流形,将过程的历史作为动态坐标嵌入。通过自洽场方法,智能体持续维护对未来路径律的预期代理,从而实现期望回报的确定性评估。该方法将随机分支转换为单次线性计算,显著降低计算复杂度与方差。我们证明该框架保持了基本收缩性质,在重尾噪声下仍能确保稳定泛化。结果表明,将强化学习建立在路径空间的拓扑特征上,能使智能体在高度波动的连续时间环境中实现主动风险管理与更优策略稳定性。
原文摘要 · Abstract (English)
This paper introduces Anticipatory Reinforcement Learning (ARL), a novel framework designed to bridge the gap between non-Markovian decision processes and classical reinforcement learning architectures, specifically under the constraint of a single observed trajectory. In environments characterised by jump-diffusions and structural breaks, traditional state-based methods often fail to capture the essential path-dependent geometry required for accurate foresight. We resolve this by lifting the state space into a signature-augmented manifold, where the history of the process is embedded as a dynamical coordinate. By utilising a self-consistent field approach, the agent maintains an anticipated proxy of the future path-law, allowing for a deterministic evaluation of expected returns. This transition from stochastic branching to a single-pass linear evaluation significantly reduces computational complexity and variance. We prove that this framework preserves fundamental contraction properties and ensures stable generalisation even in the presence of heavy-tailed noise. Our results demonstrate that by grounding reinforcement learning in the topological features of path-space, agents can achieve proactive risk management and superior policy stability in highly volatile, continuous-time environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。