用3D可操作性引导稀疏扩散策略,让双臂机器人在随机环境中精准抓取物体。
AnchorDP3: 3D Affordance Guided Sparse Diffusion Policy for Robotic Manipulation

- 通过渲染真实标签分割关键物体,提供强可操作性先验
- 98.7%成功率,极端随机条件下多任务抓取表现最优
- 无需人工示范,可自动生成部署级视觉运动策略
我们提出AnchorDP3,一种用于双臂机器人操作的扩散策略框架,在高度随机化环境中达到当前最优性能。该框架集成三项核心创新:(1) 模拟器监督语义分割,利用渲染的真实标签显式分割点云中的任务关键物体,提供强可操作性先验;(2) 任务条件特征编码器,轻量模块对每项任务处理增强点云,通过共享扩散型动作专家实现高效多任务学习;(3) 可操作性锚定关键姿态扩散与全状态监督,将密集轨迹预测改为稀疏、几何有意义的动作锚点(如预抓取、抓取姿态),直接锚定在可操作性上,大幅简化预测空间;动作专家同时预测机械臂关节角与末端执行器位姿,利用几何一致性加速收敛并提升精度。在大规模程序生成的仿真数据上训练,AnchorDP3在RoboTwin基准测试中对多样化任务实现98.7%平均成功率,即使在物体、杂乱度、桌面高度、光照和背景极端随机条件下依然稳定表现。该框架结合RoboTwin真实到模拟流水线,具备从场景和指令自动生成可部署视觉运动策略的能力,完全消除学习操作技能时的人工示范需求。
原文摘要 · Abstract (English)
We present AnchorDP3, a diffusion policy framework for dual-arm robotic manipulation that achieves state-of-the-art performance in highly randomized environments. AnchorDP3 integrates three key innovations: (1) Simulator-Supervised Semantic Segmentation, using rendered ground truth to explicitly segment task-critical objects within the point cloud, which provides strong affordance priors; (2) Task-Conditioned Feature Encoders, lightweight modules processing augmented point clouds per task, enabling efficient multi-task learning through a shared diffusion-based action expert; (3) Affordance-Anchored Keypose Diffusion with Full State Supervision, replacing dense trajectory prediction with sparse, geometrically meaningful action anchors, i.e., keyposes such as pre-grasp pose, grasp pose directly anchored to affordances, drastically simplifying the prediction space; the action expert is forced to predict both robot joint angles and end-effector poses simultaneously, which exploits geometric consistency to accelerate convergence and boost accuracy. Trained on large-scale, procedurally generated simulation data, AnchorDP3 achieves a 98.7% average success rate in the RoboTwin benchmark across diverse tasks under extreme randomization of objects, clutter, table height, lighting, and backgrounds. This framework, when integrated with the RoboTwin real-to-sim pipeline, has the potential to enable fully autonomous generation of deployable visuomotor policies from only scene and instruction, totally eliminating human demonstrations from learning manipulation skills.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。