arXiv:2509.18676cs.ROcs.SY2025-09被引 11

用3D流生成提升机器人抓取的精准度和泛化能力

3D Flow Diffusion Policy: Visuomotor Policy Learning via Generating Flow in 3D Space

  • 通过3D空间流作为中间表示,捕捉局部运动细节
  • 在50个任务中达到顶尖性能,尤其在复杂场景表现突出
  • 适合需要精细接触操作的机器人应用

学习能跨多种物体和交互动态泛化的鲁棒视觉-运动策略仍是机器人操作的核心挑战。现有方法多依赖观测到动作的直接映射,或压缩感知输入为全局或物体中心特征,常忽略精确、接触密集操作所需的局部运动线索。本文提出3D Flow Diffusion Policy(3D FDP),利用场景级3D流作为结构化中间表示,捕获细粒度局部运动线索。该方法预测采样查询点的时间轨迹,并基于这些交互感知的流来生成动作,统一整合于扩散架构中。此设计将操作锚定在局部动态,同时使策略能够推理动作对整体场景的影响。在MetaWorld基准上的大量实验表明,3D FDP在50个任务中均达到当前最优表现,尤其在中等和高难度设置下显著领先。此外,在8个真实机器人任务中验证,其在接触密集和非抓握场景中持续优于以往基线。结果表明,3D流是学习可泛化视觉-运动策略的强大结构先验,有助于构建更鲁棒、多功能的机器人操作能力。机器人演示、额外结果与代码见https://sites.google.com/view/3dfdp/home。

原文摘要 · Abstract (English)

Learning robust visuomotor policies that generalize across diverse objects and interaction dynamics remains a central challenge in robotic manipulation. Most existing approaches rely on direct observation-to-action mappings or compress perceptual inputs into global or object-centric features, which often overlook localized motion cues critical for precise and contact-rich manipulation. We present 3D Flow Diffusion Policy (3D FDP), a novel framework that leverages scene-level 3D flow as a structured intermediate representation to capture fine-grained local motion cues. Our approach predicts the temporal trajectories of sampled query points and conditions action generation on these interaction-aware flows, implemented jointly within a unified diffusion architecture. This design grounds manipulation in localized dynamics while enabling the policy to reason about broader scene-level consequences of actions. Extensive experiments on the MetaWorld benchmark show that 3D FDP achieves state-of-the-art performance across 50 tasks, particularly excelling on medium and hard settings. Beyond simulation, we validate our method on eight real-robot tasks, where it consistently outperforms prior baselines in contact-rich and non-prehensile scenarios. These results highlight 3D flow as a powerful structural prior for learning generalizable visuomotor policies, supporting the development of more robust and versatile robotic manipulation. Robot demonstrations, additional results, and code can be found at https://sites.google.com/view/3dfdp/home.

机器人操作扩散模型3D流视觉运动策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。