arXiv:2601.10471cs.LG2026-01被引 4

DeFlow解耦行为流建模与价值最大化,提升离线强化学习性能。

DeFlow: Decoupling Manifold Modeling and Value Maximization for Offline Policy Extraction

  • 分离流模型与价值优化,用轻量修正模块增强生成策略
  • 在OGBench上表现优于现有方法,实现高效离线到在线迁移
  • 无需反向传播求导,训练更稳定,适合复杂策略学习

我们提出DeFlow,一种解耦的离线强化学习框架,利用流匹配精准捕捉复杂行为流形。生成策略优化通常需对ODE求解器进行反向传播,计算成本高。我们通过在显式数据导出的信任区域内学习轻量级修正模块,避免了单步蒸馏带来的迭代生成能力损失。该方法规避了求解器微分,无需权衡损失项,确保稳定提升的同时完整保留流模型的迭代表达能力。实验表明,DeFlow在具有挑战性的OGBench基准上表现优异,并展现出高效的离线到在线适应能力。

原文摘要 · Abstract (English)

We present DeFlow, a decoupled offline RL framework that leverages flow matching to faithfully capture complex behavior manifolds. Optimizing generative policies is computationally prohibitive, typically necessitating backpropagation through ODE solvers. We address this by learning a lightweight refinement module within an explicit, data-derived trust region of the flow manifold, rather than sacrificing the iterative generation capability via single-step distillation. This way, we bypass solver differentiation and eliminate the need for balancing loss terms, ensuring stable improvement while fully preserving the flow's iterative expressivity. Empirically, DeFlow achieves superior performance on the challenging OGBench benchmark and demonstrates efficient offline-to-online adaptation.

离线强化学习流匹配策略提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。