构建166类3D物体数据集,推动6自由度姿态估计更贴近真实场景。
Omni6D: Large-Vocabulary 3D Object Dataset for Category-Level 6D Object Pose Estimation
- 涵盖166类、4688个实例,超80万张图像,覆盖多样背景与遮挡。
- 提出对称感知评估指标,系统评测现有算法在新挑战下的表现。
- 提供迁移学习方案,助力旧模型快速适配大规模类别设置。
6D物体姿态估计旨在从单张RGBD图像中确定物体的平移、旋转和尺度。近期进展已将估计任务从实例级拓展至类别级,使模型能泛化到同一类别中未见实例。然而,现有数据集如NOCS涵盖类别有限,且常忽略现实中的遮挡等挑战。为此,我们推出Omni6D,一个涵盖广泛类别与多变背景的综合性RGBD数据集,使任务更贴近真实场景。1)数据集包含166个类别、4688个调整至标准姿态的实例,以及超过80万次捕获,显著扩展了评估范围。2)我们引入对称感知评估指标,并在Omni6D上系统评测现有算法,深入揭示新挑战与洞察。3)此外,我们提出一种有效微调方法,可将先前数据集训练的模型迁移到本研究的大词汇量设定中。我们认为该工作将推动工业与学术界在通用6D姿态估计上的新突破。
原文摘要 · Abstract (English)
6D object pose estimation aims at determining an object's translation, rotation, and scale, typically from a single RGBD image. Recent advancements have expanded this estimation from instance-level to category-level, allowing models to generalize across unseen instances within the same category. However, this generalization is limited by the narrow range of categories covered by existing datasets, such as NOCS, which also tend to overlook common real-world challenges like occlusion. To tackle these challenges, we introduce Omni6D, a comprehensive RGBD dataset featuring a wide range of categories and varied backgrounds, elevating the task to a more realistic context. 1) The dataset comprises an extensive spectrum of 166 categories, 4688 instances adjusted to the canonical pose, and over 0.8 million captures, significantly broadening the scope for evaluation. 2) We introduce a symmetry-aware metric and conduct systematic benchmarks of existing algorithms on Omni6D, offering a thorough exploration of new challenges and insights. 3) Additionally, we propose an effective fine-tuning approach that adapts models from previous datasets to our extensive vocabulary setting. We believe this initiative will pave the way for new insights and substantial progress in both the industrial and academic fields, pushing forward the boundaries of general 6D pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。