提出连续时间强化学习新方法,无需离散化即可保持动作评估能力
Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control
- 基于鞅条件与动态规划,将扩散策略得分与连续Q函数梯度关联
- 在线性二次控制问题中给出理论闭式解,验证方法有效性
- 适合需要高精度连续控制的场景,如机器人运动规划
强化学习在多个领域取得显著成果,但多数方法基于离散时间。本文提出一种面向连续时间控制的新强化学习方法,其中状态-动作动力学由随机微分方程描述。不同于传统基于价值函数的方法,本工作通过鞅条件刻画连续时间Q函数,并利用动态规划原理将扩散策略得分与学习到的连续Q函数动作梯度联系起来。这一洞察催生了连续Q得分匹配(CQSM)算法,一种基于得分的策略改进方法。该方法解决了连续时间强化学习中的长期挑战:在不依赖时间离散化的情况下维持Q函数的动作评估能力。我们进一步在框架内为线性二次(LQ)控制问题提供了理论闭式解。数值实验在模拟环境中展示了所提方法的有效性,并与主流基线进行了对比。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has achieved significant success across a wide range of domains, however, most existing methods are formulated in discrete time. In this work, we introduce a novel RL method for continuous-time control, where stochastic differential equations govern state-action dynamics. Departing from traditional value function-based approaches, our key contribution is the characterization of continuous-time Q-functions via a martingale condition and the linking of diffusion policy scores to the action gradient of a learned continuous Q-function by the dynamic programming principle. This insight motivates Continuous Q-Score Matching (CQSM), a score-based policy improvement algorithm. Notably, our method addresses a long-standing challenge in continuous-time RL: preserving the action-evaluation capability of Q-functions without relying on time discretization. We further provide theoretical closed-form solutions for linear-quadratic (LQ) control problems within our framework. Numerical results in simulated environments demonstrate the effectiveness of our proposed method and compare it to popular baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。