构建了2180万张真实世界物体的9维姿态数据集,规模超现有基准百倍。
Every9D-21M: Large-Scale Real-World 9D Canonicalization of Everyday Objects

- 通过多视角几何重建与跨实例对齐,实现大规模真实图像9D姿态标注
- 涵盖700类日常物品,图像量达2180万,是此前最大数据集的100倍以上
- 适用于9D姿态基础模型训练,尤其提升在新场景下的泛化能力
从单张真实世界图像中估计日常物体的9维姿态仍具挑战,主要受限于缺乏大规模监督。现有数据集或依赖大量合成渲染,或对真实物体覆盖有限:目前最大的真实世界9D姿态数据集仅包含17,000个标注物体,涉及9个类别。本文提出Every9D-21M,一个涵盖2180万张真实世界图像的9D姿态标注数据集,来自10.9万个以物体为中心的视频,覆盖700种日常物体类别——在图像数量和类别数上均比以往真实世界9D姿态基准高出两个数量级。为实现如此规模,我们利用物体中心视频,通过多视角几何重建物体级点云,并将相似实例对齐至统一规范坐标系。仅对少量参考物体(少于0.01%的图像)进行人工标注规范姿态,并通过跨实例对齐传播至其余样本。所有传播姿态均经多视角验证。我们还引入跨类别方向规则,以建模类别级对称性,支持对称感知评估。除建立专用训练与评估划分作为9D姿态基础模型的基准外,实验表明,在Every9D-21M上训练可显著提升在ImageNet3D和PASCAL3D+上的性能,并在HANDAL上泛化能力优于ImageNet3D训练。数据与代码已公开于https://github.com/GenIntel/Every9D。
原文摘要 · Abstract (English)
Estimating the 9D pose of everyday objects from a single real-world image remains challenging. This is largely due to the lack of large-scale supervision. Most existing datasets either rely heavily on synthetic renderings or provide limited coverage of real-world objects: the largest real-world 9D pose dataset to date contains only 17K annotated objects across 9 categories. We address this gap with Every9D-21M, a dataset of 9D pose annotations for 21.8M real-world images from 109K object- centric videos spanning 700 everyday object categories - two orders of magnitude larger than prior real-world 9D pose benchmarks in both image and category count. To achieve this scale, we leverage object-centric videos by reconstructing object- level point clouds via multi-view geometry and aligning similar instances into a shared canonical coordinate frame. Canonical poses are manually annotated for only a small set of reference objects (fewer than 0.01% of all images) and propagated to the remaining instances via cross-instance alignment. All propagated canonical poses are then verified from multiple viewpoints. We further introduce cross-category orientation rules that induce category-level symmetries, enabling symmetry-aware evaluation. Beyond establishing dedicated training and evaluation splits as a benchmark for 9D pose foundation models, we show that training on Every9D-21M improves performance on ImageNet3D and PASCAL3D+, and generalizes to HANDAL substantially better than training on ImageNet3D. Data and code are available at https://github.com/GenIntel/Every9D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。