构建6D变形物体姿态数据集,揭示现有方法在形变下的性能崩溃
Exploring 6D Object Pose Estimation with Deformation

- 构建包含26类物体的高保真3D扫描数据集,每类有1个标准态和3个变形态
- 在133K帧上生成66.5万条姿态标注,验证了形变导致性能显著下降
- 适合做物体姿态估计、形变建模与实际场景鲁棒性研究的团队参考
我们提出DeSOPE,一个面向6自由度变形物体的大规模数据集。多数6D物体姿态方法假设物体为刚体或可动结构,但在现实中,物体因磨损、撞击或形变会偏离原始形态。为此,我们构建了包含26种常见物体类别、每类具有1个标准状态和3个形变状态的高保真3D扫描数据集,并对每个扫描进行精确的3D配准至标准网格。此外,该数据集还包含跨多样化场景的13.3万帧RGB-D数据,以及通过半自动流程生成的66.5万条姿态标注。标注流程包括:首先人工标注2D掩码,再用物体姿态方法获取初始姿态,随后通过物体级SLAM系统优化,最后进行人工核查确认。我们评估多种物体姿态方法,发现随着形变程度增加,性能急剧下降,表明在实际应用中应对形变具备更强鲁棒性至关重要。
原文摘要 · Abstract (English)
We present DeSOPE, a large-scale dataset for 6DoF deformed objects. Most 6D object pose methods assume rigid or articulated objects, an assumption that fails in practice as objects deviate from their canonical shapes due to wear, impact, or deformation. To model this, we introduce the DeSOPE dataset, which features high-fidelity 3D scans of 26 common object categories, each captured in one canonical state and three deformed configurations, with accurate 3D registration to the canonical mesh. Additionally, it features an RGB-D dataset with 133K frames across diverse scenarios and 665K pose annotations produced via a semi-automatic pipeline. We begin by annotating 2D masks for each instance, then compute initial poses using an object pose method, refine them through an object-level SLAM system, and finally perform manual verification to produce the final annotations. We evaluate several object pose methods and find that performance drops sharply with increasing deformation, suggesting that robust handling of such deformations is critical for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。