用最优传输强化学习,让机器人用少量试错快速优化和加速流式策略。
Optimal Transport Q-Learning for Flow Policy Steering and Acceleration

- 基于优势加权最优传输匹配,用少量交互数据微调流式策略。
- 仅需50-60次交互,成功率从36%提至86%,推理步数减少70%。
- 适合资源有限的机器人场景,尤其适合视觉语言动作模型后训练。
扩散和流式策略在机器人应用中表现出色,能精准捕捉多模态轨迹分布,尤其在视觉-语言-动作(VLA)模型中表现突出。然而,高质量策略还需快速推理与优质示范,这些往往难以获取,导致策略性能不佳且在分布外情形下易失效。本文提出最优传输Q-learning(OTQL),一种基于强化学习的后训练方法,利用机器人自身经验微调并加速流式策略。OTQL通过优势加权条件最优传输流匹配实现高效微调,仅需50-60次交互即可完成,无需昂贵的仿真或真实世界中的知识蒸馏。实验表明,该方法使单任务策略平均成功率从36%提升至86%,预训练VLA策略从38%提升至76%,同时将每动作生成的推理步数减少70%。
原文摘要 · Abstract (English)
Diffusion and flow policies have recently demonstrated remarkable performance in robotic applications by accurately capturing multimodal robot trajectory distributions, especially in the context of vision language action (VLA) models. However, high quality policy performance also requires fast inference and high quality demonstrations, which are often hard to get. Lack of these leads to suboptimal policy behaviors and failure under distribution shifts. In this work we address the problem of fine-tuning and accelerating suboptimal flow-based policies using the robot's experience through RL post-training. We introduce Optimal Transport Q-Learning (OTQL), a new method for finetuning flow policies using advantage weighted conditional optimal transport flow matching. OTQL can finetune and accelerate flows with an interaction budget of 50-60 episodes while avoiding computationally expensive distillation in simulation and real-world robot tasks. Our results show that OTQL post-trains flow policies using the robot's own experience, increasing average success percentage of single-task policies from 36% to 86% and of a pre-trained VLA from 38% to 76% while reducing the number of inference steps per action generation by 70%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。