提出一步生成的视觉语言动作模型,速度提升超80倍
Mean-Flow based One-Step Vision-Language-Action
- 用均值流方法替代传统迭代采样,实现单步动作生成
- 实测生成速度比SmolVLA快8.7倍,比Diffusion Policy快83.9倍
- 适合对实时性要求高的机器人精细操作场景
基于流匹配的视觉语言动作(VLA)框架在生成高频动作片段方面表现出显著优势,尤其适用于高灵巧性机器人操作任务。然而,其实际应用受限于生成延迟过长,根源在于固有的迭代采样需求和架构限制。为此,我们提出基于均值流的一步式VLA方法,通过解决动作生成过程中的噪声问题,消除了传统流匹配方法固有的一致性约束,显著提升生成效率,实现单步生成。真实机器人实验表明,所提方法的生成速度分别达到SmolVLA的8.7倍和Diffusion Policy的83.9倍。结果表明,该方法具有作为VLA驱动机器人操作高效骨干的巨大潜力。
原文摘要 · Abstract (English)
Recent advances in FlowMatching-based Vision-Language-Action (VLA) frameworks have demonstrated remarkable advantages in generating high-frequency action chunks, particularly for highly dexterous robotic manipulation tasks. Despite these notable achievements, their practical applications are constrained by prolonged generation latency, which stems from inherent iterative sampling requirements and architectural limitations. To address this critical bottleneck, we propose a Mean-Flow based One-Step VLA approach. Specifically, we resolve the noise-induced issues in the action generation process, thereby eliminating the consistency constraints inherent to conventional Flow-Matching methods. This significantly enhances generation efficiency and enables one-step action generation. Real-world robotic experiments show that the generation speed of the proposed Mean-Flow based One-Step VLA is 8.7 times and 83.9 times faster than that of SmolVLA and Diffusion Policy, respectively. These results elucidate its great potential as a high-efficiency backbone for VLA-based robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。