arXiv:2602.13810cs.LGcs.AI2026-02被引 9

提出一种一步生成动作的高效策略,提升机器人抓取任务表现

Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation

  • 用均速度场建模策略,实现单步快速采样
  • 引入瞬时速度约束,提升策略表达能力与学习精度
  • 在多个机器人任务中性能领先,训练推理更快

在强化学习中,学习具有表现力且高效的策略函数是一条有前景的方向。尽管基于流的策略最近在以快速确定性采样方式建模复杂动作分布方面表现有效,但仍面临表达力与计算开销之间的权衡,通常由流步数控制。本文提出均速度策略(MVP),一种新的生成式策略函数,通过建模均速度场实现最快的单步动作生成。为确保高表达力,在训练过程中对均速度场引入瞬时速度约束(IVC)。我们从理论上证明,该设计明确作为关键边界条件,从而提升学习精度并增强策略表达力。实验表明,MVP在Robomimic和OGBench的多个挑战性机器人操作任务中达到当前最优成功率,同时在训练和推理速度上显著优于现有基于流的策略基线。

原文摘要 · Abstract (English)

Learning expressive and efficient policy functions is a promising direction in reinforcement learning (RL). While flow-based policies have recently proven effective in modeling complex action distributions with a fast deterministic sampling process, they still face a trade-off between expressiveness and computational burden, which is typically controlled by the number of flow steps. In this work, we propose mean velocity policy (MVP), a new generative policy function that models the mean velocity field to achieve the fastest one-step action generation. To ensure its high expressiveness, an instantaneous velocity constraint (IVC) is introduced on the mean velocity field during training. We theoretically prove that this design explicitly serves as a crucial boundary condition, thereby improving learning accuracy and enhancing policy expressiveness. Empirically, our MVP achieves state-of-the-art success rates across several challenging robotic manipulation tasks from Robomimic and OGBench. It also delivers substantial improvements in training and inference speed over existing flow-based policy baselines.

强化学习策略生成机器人控制流模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。