用分布式强化学习解决混沌系统中的高方差问题
On Distributional Reinforcement Learning in Chaotic Dynamical Systems

- 基于1-Wasserstein距离优化回报分布,使目标更平滑
- 在混沌系统中,分布级贝尔曼目标方差显著降低
- 适合研究气候、流体等混沌动力系统的学者
混沌动力系统对强化学习构成根本挑战:对初始条件的指数敏感性导致自举目标方差过高,梯度更新条件不良。混沌动力系统广泛存在于流体流动、气候系统及多智能体系统中,可靠学习至关重要。标准RL方法通过标量值函数优化期望回报,隐式平均发散轨迹,将轨迹级不稳定性与学习目标纠缠。我们证明,在温和的统计稳定性假设下,回报分布在1-Wasserstein度量下演化比单条轨迹更规则,从而产生更平滑的分布式贝尔曼目标。通过与该度量层级结构对齐,分布式RL实现更良态的学习。本文为分布式方法在混沌系统中的优势提供了原则性解释,并揭示了混沌下RL目标的几何特性。
原文摘要 · Abstract (English)
Chaotic dynamical systems pose a fundamental challenge for Reinforcement Learning (RL): exponential sensitivity to initial conditions induces high-variance bootstrap targets and poorly conditioned gradient updates. Chaotic dynamics arise across scientific and engineering domains, from fluid flows and climate systems to multi-agent systems, where reliable learning is highly desirable. Standard RL methods optimise expected returns through scalar value functions, implicitly averaging over diverging trajectories and entangling trajectory level instability with the learning objective. We show that under mild statistical stability assumptions, the return distribution evolves more regularly than individual trajectories when measured under the $1$-Wasserstein metric, yielding a smoother distributional Bellman objective. By aligning optimisation with this measure level structure, distributional RL provides better conditioned learning. We offer a principled explanation for the advantages of distributional methods in chaotic systems and the geometries of RL objectives under chaos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。