arXiv:2603.12480cs.ROcs.AI2026-03被引 6

提出单步生成动作的机器人控制方法,速度提升100倍且精度更高。

One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies

  • 通过自蒸馏构建无需预训练教师的单步动作生成框架
  • 56个仿真任务中单步性能超越100步扩散模型,提速超100倍
  • 适合对实时性要求高的机器人精准操控场景

生成流和扩散模型能提供高精度机器人策略所需的连续、多模态动作分布。然而,其依赖迭代采样导致严重推理延迟,降低控制频率并影响时间敏感操作的表现。为此,我们提出一种从零开始的自蒸馏框架——单步流策略(OFP),实现高保真、单步动作生成。OFP统一了自一致性损失以确保时间区间间一致的迁移,以及自引导正则化以使预测聚焦于高密度专家模式。此外,热启动机制利用动作的时间相关性最小化生成传输距离。在56个多样化的仿真操控任务中,单步OFP达到领先水平,性能超越100步扩散和流模型,同时动作生成速度提升超过100倍。我们将OFP集成到RoboTwin 2.0的π_{0.5}模型中,单步OFP表现优于原10步策略。结果表明,OFP是高精度与低延迟机器人控制的实用且可扩展解决方案。

原文摘要 · Abstract (English)

Generative flow and diffusion models provide the continuous, multimodal action distributions needed for high-precision robotic policies. However, their reliance on iterative sampling introduces severe inference latency, degrading control frequency and harming performance in time-sensitive manipulation. To address this problem, we propose the One-Step Flow Policy (OFP), a from-scratch self-distillation framework for high-fidelity, single-step action generation without a pre-trained teacher. OFP unifies a self-consistency loss to enforce coherent transport across time intervals, and a self-guided regularization to sharpen predictions toward high-density expert modes. In addition, a warm-start mechanism leverages temporal action correlations to minimize the generative transport distance. Evaluations across 56 diverse simulated manipulation tasks demonstrate that a one-step OFP achieves state-of-the-art results, outperforming 100-step diffusion and flow policies while accelerating action generation by over $100\times$. We further integrate OFP into the $π_{0.5}$ model on RoboTwin 2.0, where one-step OFP surpasses the original 10-step policy. These results establish OFP as a practical, scalable solution for highly accurate and low-latency robot control.

机器人控制单步生成自蒸馏低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。