arXiv:2601.20701cs.RO2026-01被引 6

提出一步生成策略优化方法,实现机器人实时控制的高速推理。

One Step Is Enough: Dispersive MeanFlow Policy Optimization

  • 基于均值流设计单步推断,无需多步采样和知识蒸馏
  • 在多个基准上性能媲美或超越多步基线,推理速度提升5-20倍
  • 适合对实时性要求高的机器人控制场景,已实现在真实机械臂部署

实时机器人控制需要快速动作生成。然而,现有的基于扩散和流匹配的生成策略依赖多步采样,从根本上限制了其在时间敏感场景中的应用。我们提出分散式均值流策略优化(DMPO),一个统一框架,通过三个关键组件实现真正的单步生成:均值流用于数学推导的单步推理,无需知识蒸馏;分散正则化防止表示崩溃;强化学习微调使性能超越专家示范。在RoboMimic操控和OpenAI Gym运动基准上的实验表明,其性能与多步基线相当或更优。结合轻量模型架构与三项算法组件的协同作用,DMPO满足实时控制需求(>120Hz),在高性能GPU上达到数百赫兹,推理速度提升5-20倍。在Franka Emika Panda机器人上的物理部署验证了其实际可行性。

原文摘要 · Abstract (English)

Real-time robotic control demands fast action generation. However, existing generative policies based on diffusion and flow matching require multi-step sampling, fundamentally limiting deployment in time-critical scenarios. We propose Dispersive MeanFlow Policy Optimization (DMPO), a unified framework that enables true one-step generation through three key components: MeanFlow for mathematically-derived single-step inference without knowledge distillation, dispersive regularization to prevent representation collapse, and reinforcement learning (RL) fine-tuning to surpass expert demonstrations. Experiments across RoboMimic manipulation and OpenAI Gym locomotion benchmarks demonstrate competitive or superior performance compared to multi-step baselines. With our lightweight model architecture and the three key algorithmic components working in synergy, DMPO exceeds real-time control requirements (>120Hz) with 5-20x inference speedup, reaching hundreds of Hertz on high-performance GPUs. Physical deployment on a Franka-Emika-Panda robot validates real-world applicability.

机器人控制单步生成强化学习实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。