arXiv:2508.11002cs.RO2025-08被引 22

3D FlowMatch Actor实现单双臂机器人操作的高效统一控制

3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation

  • 结合流匹配与3D视觉表征,直接预测末端执行器轨迹
  • 在双臂任务上超越前人41.4%,训练推理速度提升30倍以上
  • 适合需要高精度、低延迟机器人控制的应用场景

我们提出3D FlowMatch Actor(3DFA),一种用于机器人操作的3D策略架构,融合流匹配轨迹预测与3D预训练视觉场景表征以实现从示范学习。3DFA在动作去噪过程中引入动作与视觉标记间的3D相对注意力,基于先前的3D扩散单臂策略工作。通过流匹配与系统级及架构优化的结合,3DFA在不牺牲性能的前提下,实现比以往3D扩散策略快30倍以上的训练与推理速度。在双臂任务基准PerAct2上,其表现超越次优方法41.4%的绝对差距,创下新纪录。在大量真实世界评估中,它优于具有最多达1000倍参数和更长预训练时间的强基线模型。在单臂设置下,其在74个RLBench任务上直接预测密集末端执行器轨迹,无需运动规划,创下新纪录。全面消融实验验证了设计选择对策略有效性和效率的重要性。

原文摘要 · Abstract (English)

We present 3D FlowMatch Actor (3DFA), a 3D policy architecture for robot manipulation that combines flow matching for trajectory prediction with 3D pretrained visual scene representations for learning from demonstration. 3DFA leverages 3D relative attention between action and visual tokens during action denoising, building on prior work in 3D diffusion-based single-arm policy learning. Through a combination of flow matching and targeted system-level and architectural optimizations, 3DFA achieves over 30x faster training and inference than previous 3D diffusion-based policies, without sacrificing performance. On the bimanual PerAct2 benchmark, it establishes a new state of the art, outperforming the next-best method by an absolute margin of 41.4%. In extensive real-world evaluations, it surpasses strong baselines with up to 1000x more parameters and significantly more pretraining. In unimanual settings, it sets a new state of the art on 74 RLBench tasks by directly predicting dense end-effector trajectories, eliminating the need for motion planning. Comprehensive ablation studies underscore the importance of our design choices for both policy effectiveness and efficiency.

机器人控制3D扩散模型端到端策略多臂操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。