arXiv:2602.14193cs.ROcs.CV2026-02中稿 · ICLR被引 4

提出3D功能部件感知特征场,提升机器人操作不同关节物体的泛化能力。

Learning Part-Aware Dense 3D Feature Field for Generalizable Articulated Object Manipulation

  • 构建基于3D点云的部件感知特征场,通过对比学习训练
  • 在模拟与真实任务中,性能超越CLIP、DINOv2等2D/3D基线模型
  • 适用于模仿学习、对应点匹配与分割,可作为通用机器人基础框架

关节物体操作对各类实际机器人任务至关重要,但跨多样物体的泛化仍是重大挑战。关键在于理解功能部件(如门把手、旋钮),它们指示了不同类别和形状物体的可操作位置与方式。以往工作尝试通过引入基础特征实现泛化,但这些特征多为2D且未专门考虑功能部件。将2D特征映射至几何丰富的3D空间时,存在运行时间长、多视角不一致、空间分辨率低等问题。为此,我们提出部件感知3D特征场(PA3FF),一种具备部件意识的新型密集3D特征表示,用于通用关节物体操作。PA3FF通过大规模标注数据集中的3D部件提案进行训练,采用对比学习范式。输入点云后,其以前馈方式预测连续3D特征场,特征间距离反映功能部件的接近程度:特征相似的点更可能属于同一部件。基于此特征,我们设计了部件感知扩散策略(PADP),一种面向样本效率与泛化的模仿学习框架。我们在多个模拟与真实任务上评估了PADP,结果表明,相较于CLIP、DINOv2、Grounded-SAM等2D/3D表示,PA3FF在各类操作场景中均表现更优。此外,PA3FF还可支持多种下游任务,包括对应关系学习与分割,具备强通用性。

原文摘要 · Abstract (English)

Articulated object manipulation is essential for various real-world robotic tasks, yet generalizing across diverse objects remains a major challenge. A key to generalization lies in understanding functional parts (e.g., door handles and knobs), which indicate where and how to manipulate across diverse object categories and shapes. Previous works attempted to achieve generalization by introducing foundation features, while these features are mostly 2D-based and do not specifically consider functional parts. When lifting these 2D features to geometry-profound 3D space, challenges arise, such as long runtimes, multi-view inconsistencies, and low spatial resolution with insufficient geometric information. To address these issues, we propose Part-Aware 3D Feature Field (PA3FF), a novel dense 3D feature with part awareness for generalizable articulated object manipulation. PA3FF is trained by 3D part proposals from a large-scale labeled dataset, via a contrastive learning formulation. Given point clouds as input, PA3FF predicts a continuous 3D feature field in a feedforward manner, where the distance between point features reflects the proximity of functional parts: points with similar features are more likely to belong to the same part. Building on this feature, we introduce the Part-Aware Diffusion Policy (PADP), an imitation learning framework aimed at enhancing sample efficiency and generalization for robotic manipulation. We evaluate PADP on several simulated and real-world tasks, demonstrating that PA3FF consistently outperforms a range of 2D and 3D representations in manipulation scenarios, including CLIP, DINOv2, and Grounded-SAM. Beyond imitation learning, PA3FF enables diverse downstream methods, including correspondence learning and segmentation tasks, making it a versatile foundation for robotic manipulation. Project page: https://pa3ff.github.io

3D特征场关节物体机器人操作模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。