arXiv:2510.15388cs.LG2025-10被引 1

提出可在线稳定微调扩散策略的新框架,解决演示学习后分布偏移问题。

Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning

  • 将流匹配分解为小步更新,每步对应最优传输中的JKO变分更新
  • 在多个机器人控制任务上实现更稳定、更快的在线适应,计算开销降低30%以上
  • 适合需要持续学习与快速响应的机器人系统,尤其擅长复杂动作微调

尽管基于流/扩散策略的行为克隆能从示范中学习复杂技能,但仍易受分布偏移影响;而标准强化学习方法因迭代推理过程和现有方案局限,难以有效微调此类模型。本文提出步进流策略(SWFP)框架,其核心思想是:通过固定步长欧拉法离散化流匹配推理过程,天然契合最优传输中的变分Jordan-Kinderlehrer-Otto(JKO)原理。SWFP将全局流分解为一系列相邻分布间的微小增量变换,每一步对应一次JKO更新,通过熵正则化使策略变化贴近前一迭代点,保障在线适应的稳定性。该分解带来高效算法:通过级联小规模流模块微调预训练流模型,实现子模型训练更简单快速、计算与内存成本显著降低,并具备基于Wasserstein信任区域的可证明稳定性。大量实验表明,SWFP在多种机器人控制基准测试中展现出更强稳定性、更高效率和更优适应性能。

原文摘要 · Abstract (English)

While behavior cloning with flow/diffusion policies excels at learning complex skills from demonstrations, it remains vulnerable to distributional shift, and standard RL methods struggle to fine-tune these models due to their iterative inference process and the limitations of existing workarounds. In this work, we introduce the Stepwise Flow Policy (SWFP) framework, founded on the key insight that discretizing the flow matching inference process via a fixed-step Euler scheme inherently aligns it with the variational Jordan-Kinderlehrer-Otto (JKO) principle from optimal transport. SWFP decomposes the global flow into a sequence of small, incremental transformations between proximate distributions. Each step corresponds to a JKO update, regularizing policy changes to stay near the previous iterate and ensuring stable online adaptation with entropic regularization. This decomposition yields an efficient algorithm that fine-tunes pre-trained flows via a cascade of small flow blocks, offering significant advantages: simpler/faster training of sub-models, reduced computational/memory costs, and provable stability grounded in Wasserstein trust regions. Comprehensive experiments demonstrate SWFP's enhanced stability, efficiency, and superior adaptation performance across diverse robotic control benchmarks.

强化学习扩散模型在线学习机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。