arXiv:2509.01819cs.RO2025-09被引 43

用一致性流训练实现高精度机器人操作,1-2步生成复杂动作

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training

  • 基于流匹配与一致性训练,1-2步生成高维精细动作
  • 在真实场景中成功率接近翻倍,适配单臂、双臂和人形机器人
  • 支持视觉、语言、本体感知多模态输入,泛化能力强

本文提出ManiFlow,一种通用机器人操作的视觉-运动模仿学习策略,能根据多样化的视觉、语言和本体感知输入生成高精度、高维度的动作。通过流匹配结合一致性训练,仅需1-2次推理步骤即可实现高质量灵巧动作生成。为高效处理多模态输入,提出DiT-X:一种带有自适应交叉注意力和AdaLN-Zero条件化的扩散变换器架构,支持动作标记与多模态观测间的细粒度特征交互。ManiFlow在多个仿真基准上表现一致提升,在真实世界任务中成功率近乎翻倍,覆盖单臂、双臂及人形机器人系统,且对新物体和背景变化具有强鲁棒性与泛化能力。大规模数据集下展现显著可扩展性。

原文摘要 · Abstract (English)

This paper introduces ManiFlow, a visuomotor imitation learning policy for general robot manipulation that generates precise, high-dimensional actions conditioned on diverse visual, language and proprioceptive inputs. We leverage flow matching with consistency training to enable high-quality dexterous action generation in just 1-2 inference steps. To handle diverse input modalities efficiently, we propose DiT-X, a diffusion transformer architecture with adaptive cross-attention and AdaLN-Zero conditioning that enables fine-grained feature interactions between action tokens and multi-modal observations. ManiFlow demonstrates consistent improvements across diverse simulation benchmarks and nearly doubles success rates on real-world tasks across single-arm, bimanual, and humanoid robot setups with increasing dexterity. The extensive evaluation further demonstrates the strong robustness and generalizability of ManiFlow to novel objects and background changes, and highlights its strong scaling capability with larger-scale datasets. Our website: maniflow-policy.github.io.

机器人操作扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。