用虚拟现实高效标注可动部件结构,提升机器人抓取泛化能力。
Revisiting Articulated Parts Perception in Robot Manipulation

- 提出几何主结构GPS抽象,平衡标注效率与精度
- 仅用1分钟/物体标注41K帧数据,支持9种物体270种初始状态
- 无需领域微调,73%成功率,适合工业场景快速部署
我们日常接触的物体多具可动部件(如盒子、把手、门)。准确且通用的可动部件感知对提升机器人操作能力至关重要。现有方法分两类:基于姿态的表示需大量人工标注,成本高;基于功能的估计虽免人工但数据质量差。本文提出几何主结构(GPS)新表征,抽象部件几何结构,在可扩展性与质量间取得平衡。结合便携式虚拟现实设备,单物体序列标注仅需1分钟。该直接人工标注方式数据质量优于估算结果。通过此系统,我们收集了234个物体、6类部件共41,000帧数据,训练出仅需单张RGB-D图像输入的通用GPS模型。在实际操作中,采用基于GPS预测的启发式策略,无需领域微调即实现73%成功率,覆盖9个物体的270种初始状态。代码、数据及工具已公开。
原文摘要 · Abstract (English)
We are surrounded by various objects with movable, articulated parts, e.g., box, handle, door. An accurate and generalizable perception of articulated parts is essential to enhance robotic manipulation capabilities. Building on this need, recent efforts in articulated parts perception have followed two main directions: One line of work uses pose-based representation, which requires high manual cost; in parallel, affordance-based methods extract future object motion from point tracking without additional manual efforts, but suffer from low-quality data. In this paper, we propose a new representation of articulated parts, Geometric Primary Structure (GPS), an abstraction of the part geometry structure to balance scalability and quality. For efficient and scalable data collection, GPS is integrated with a portable Virtual Reality (VR) device and requires only one minute to annotate one object sequence. This direct human annotation provides higher quality than the estimated affordance. With this efficient VR-GPS system, we collect 41K frames for 234 objects across six part classes, and train a generalizable GPS model with a single RGB-D object image as input. For object manipulation, we deploy a heuristic policy based on GPS prediction. Without any in-domain fine-tuning, our method achieves an 73% success rate, covering 270 initial states for 9 objects. Our code, data and reusable tool are available at https://enlighten0707.github.io/gps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。