arXiv:2505.14821cs.LGcs.AI2025-05中稿 · UAI 2025被引 2

提出首个具理论保障的连续时间强化学习算法,高效学习最优策略。

Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation

  • 用乐观置信集设计模型基于的高效算法
  • 仅需N次测量即可达近优策略,误差随√N下降
  • 大幅减少策略更新与轨迹采样,适合资源受限场景

连续时间强化学习(CTRL)为连续演化环境中的序贯决策提供了严谨框架。尽管其在实践中表现良好,但针对一般函数逼近设置的理论理解仍有限。本文提出一种模型基的CTRL算法,实现样本与计算双重高效。通过乐观置信集,首次建立带一般函数逼近的CTRL样本复杂度保证:使用N次观测,可实现近优策略,子优性差距为$ ilde{O}( ext{sqrt}(d_{ ext{R}} + d_{ ext{F}})N^{-1/2})$,其中$d_{ ext{R}}$和$d_{ ext{F}}$分别为奖励函数与动态函数的分布埃尔德维维数,刻画了函数逼近的复杂度。此外,引入结构化策略更新与替代测量策略,显著减少策略更新次数与轨迹采样量,同时保持优异样本效率。实验在连续控制任务与扩散模型微调上验证了算法有效性,性能相当但策略更新与滚落次数大幅降低。

原文摘要 · Abstract (English)

Continuous-time reinforcement learning (CTRL) provides a principled framework for sequential decision-making in environments where interactions evolve continuously over time. Despite its empirical success, the theoretical understanding of CTRL remains limited, especially in settings with general function approximation. In this work, we propose a model-based CTRL algorithm that achieves both sample and computational efficiency. Our approach leverages optimism-based confidence sets to establish the first sample complexity guarantee for CTRL with general function approximation, showing that a near-optimal policy can be learned with a suboptimality gap of $\tilde{O}(\sqrt{d_{\mathcal{R}} + d_{\mathcal{F}}}N^{-1/2})$ using $N$ measurements, where $d_{\mathcal{R}}$ and $d_{\mathcal{F}}$ denote the distributional Eluder dimensions of the reward and dynamic functions, respectively, capturing the complexity of general function approximation in reinforcement learning. Moreover, we introduce structured policy updates and an alternative measurement strategy that significantly reduce the number of policy updates and rollouts while maintaining competitive sample efficiency. We implemented experiments to backup our proposed algorithms on continuous control tasks and diffusion model fine-tuning, demonstrating comparable performance with significantly fewer policy updates and rollouts.

强化学习连续时间函数逼近高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。