arXiv:2412.01398cs.CVcs.RO2024-12ICCV被引 7

构建首个统一标注的3D可动物体数据集,支持场景理解与机器人操作。

Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description

  • 提出统一框架USDNet,同时预测物体部件分割与运动属性。
  • 在280个场景中提供8类标注,覆盖部件与运动信息。
  • 适合研究具身智能、混合现实与机器人交互的开发者使用。

3D场景理解是计算机视觉中的长期挑战,也是混合现实、可穿戴计算和具身AI的关键组成部分。解决这些应用需求需要兼顾场景中心、物体中心及交互中心的能力。尽管已有大量数据集和算法处理前两者,但对可动、可交互物体的理解仍不足。本文提出:(1) Articulate3D,一个由专家精心构建的3D数据集,包含280个室内场景的高质量人工标注,提供8种类型标注,涵盖部件与详细运动信息,采用标准化场景表示格式,适用于大规模3D内容创建、交换及模拟环境无缝集成;(2) USDNet,一种新型统一框架,可同时预测部件分割与完整运动属性。我们在Articulate3D及两个现有数据集上评估USDNet,验证了统一密集预测方法的优势。通过跨数据集与跨域评估,凸显Articulate3D价值,并展示其在下游任务如基于大模型提示的场景编辑与可动物体操作机器人策略训练中的适用性。数据集、基准与代码已开源。

原文摘要 · Abstract (English)

3D scene understanding is a long-standing challenge in computer vision and a key component in enabling mixed reality, wearable computing, and embodied AI. Providing a solution to these applications requires a multifaceted approach that covers scene-centric, object-centric, as well as interaction-centric capabilities. While there exist numerous datasets and algorithms approaching the former two problems, the task of understanding interactable and articulated objects is underrepresented and only partly covered in the research field. In this work, we address this shortcoming by introducing: (1) Articulate3D, an expertly curated 3D dataset featuring high-quality manual annotations on 280 indoor scenes. Articulate3D provides 8 types of annotations for articulated objects, covering parts and detailed motion information, all stored in a standardized scene representation format designed for scalable 3D content creation, exchange and seamless integration into simulation environments. (2) USDNet, a novel unified framework capable of simultaneously predicting part segmentation along with a full specification of motion attributes for articulated objects. We evaluate USDNet on Articulate3D as well as two existing datasets, demonstrating the advantage of our unified dense prediction approach. Furthermore, we highlight the value of Articulate3D through cross-dataset and cross-domain evaluations and showcase its applicability in downstream tasks such as scene editing through LLM prompting and robotic policy training for articulated object manipulation. We provide open access to our dataset, benchmark, and method's source code.

3D理解可动物体具身智能数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。