arXiv:2605.12416cs.LG2026-05

提出快速生成动作的流映射策略,提升机器人控制效率。

Aligning Flow Map Policies with Optimal Q-Guidance

论文配图:Aligning Flow Map Policies with Optimal Q-Guidance
图 1 · 摘自论文原文
  • 设计流映射策略,可一步跳转生成动作,大幅降低推理延迟。
  • 在12个任务上平均成功率提升21.3%,优于现有方法。
  • 适合需要高效在线适应的机器人控制场景。

基于扩散和流匹配等表达性强的生成策略,适用于具有高度多模态动作分布的复杂控制问题。但其高表达性带来显著推理成本:每次动作生成需多次模拟生成过程,导致序列决策中延迟累积。本文提出流映射策略(flow map policies),通过学习在现有流基策略的生成动力学中执行任意大小跳跃(包括一步跳跃),实现快速动作生成。针对离线到在线强化学习,将在线适配建模为信任域优化问题,在提升评价器Q值的同时保持与离线策略接近。理论上推导出最优的闭式学习目标——FLOW MAP Q-GUIDANCE(FMQ),用于在评价值引导的信任域约束下适配离线流映射策略。进一步提出Q-GUIDED BEAM SEARCH(QGBS),结合重噪声与束搜索,支持推理时迭代优化。在OGBench和RoboMimic的12个挑战性机器人操作与运动任务中,FMQ在离线到在线强化学习中达到当前最优表现,平均成功率相比此前一步策略MVP提升21.3%。

原文摘要 · Abstract (English)

Generative policies based on expressive model classes, such as diffusion and flow matching, are well-suited to complex control problems with highly multimodal action distributions. Their expressivity, however, comes at a significant inference cost: generating each action typically requires simulating many steps of the generative process, compounding latency across sequential decision-making rollouts. We introduce flow map policies, a novel class of generative policies designed for fast action generation by learning to take arbitrary-size jumps including one-step jumps-across the generative dynamics of existing flow-based policies. We instantiate flow map policies for offline-to-online reinforcement learning (RL) and formulate online adaptation as a trust-region optimization problem that improves the critic's Q-value while remaining close to the offline policy. We theoretically derive FLOW MAP Q-GUIDANCE (FMQ), a principled closed-form learning target that is optimal for adapting offline flow map policies under a critic-guided trust-region constraint. We further introduce Q-GUIDED BEAM SEARCH (QGBS), a stochastic flow-map sampler that combines renoising with beam search to enable iterative inference-time refinement. Across 12 challenging robotic manipulation and locomotion tasks from OGBench and RoboMimic, FMQ achieves state-of-the-art performance in offline-to-online RL, outperforming the previous one-step policy MVP by a relative improvement of 21.3% on the average success rate.

强化学习流映射机器人控制快速生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。