arXiv:2511.06816cs.LGcs.AI2025-11被引 1

用流匹配生成高质量轨迹,提升在线强化学习的样本效率与泛化能力。

Controllable Flow Matching for Online Reinforcement Learning

  • 基于条件流匹配直接建模高回报轨迹分布,不依赖环境转移函数。
  • 通过最小化控制能量实现最优轨迹采样,在MuJoCo上性能超越传统动力学模型。
  • 适合需要高效探索和跨任务泛化的在线强化学习场景。

基于模型的强化学习(MBRL)通常依赖环境动态建模以提高数据效率。然而,长期轨迹中模型误差累积导致建模不稳定。为此,我们提出CtrlFlow,一种基于条件流匹配(CFM)的轨迹级合成方法,直接建模从初始状态到高回报终止状态的轨迹分布,无需显式建模环境转移函数。该方法通过最小化由非线性可控性格拉姆矩阵决定的控制能量,确保最优轨迹采样;生成的多样化轨迹数据显著提升了策略学习的鲁棒性和跨任务泛化能力。在在线设置下,CtrlFlow在常见的MuJoCo基准任务上表现优于动力学模型,并展现出比标准MBRL方法更优的样本效率。

原文摘要 · Abstract (English)

Model-based reinforcement learning (MBRL) typically relies on modeling environment dynamics for data efficiency. However, due to the accumulation of model errors over long-horizon rollouts, such methods often face challenges in maintaining modeling stability. To address this, we propose CtrlFlow, a trajectory-level synthetic method using conditional flow matching (CFM), which directly modeling the distribution of trajectories from initial states to high-return terminal states without explicitly modeling the environment transition function. Our method ensures optimal trajectory sampling by minimizing the control energy governed by the non-linear Controllability Gramian Matrix, while the generated diverse trajectory data significantly enhances the robustness and cross-task generalization of policy learning. In online settings, CtrlFlow demonstrates the better performance on common MuJoCo benchmark tasks than dynamics models and achieves superior sample efficiency compared to standard MBRL methods.

强化学习流匹配在线学习轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。