从单图对不精确3D模型进行零样本精准对齐,无需标注
Zero-shot Inexact CAD Model Alignment from a Single Image
- 基于基础特征构建多视角一致新空间,自监督消除对称歧义
- 在ScanNet25k上比最先进弱监督方法高4.3%平均对齐精度
- 首次实现零样本泛化,在20类新物体上超越有监督基线
从单张图像推断3D场景结构的一种实用方法是:从数据库中检索一个匹配的3D模型并将其与图像中的物体对齐。现有方法依赖于带姿态标注的监督训练,仅适用于有限类别。为此,我们提出一种无需姿态标注的弱监督9自由度对齐方法,可泛化至未见类别。该方法基于基础特征构建新型特征空间,通过自监督三元组损失确保多视角一致性,并克服基础特征固有的对称性歧义。此外,我们引入纹理无关的姿态精修技术,在归一化物体坐标下实现密集对齐,利用增强后的特征空间进行估计。我们在真实世界数据集ScanNet25k上进行广泛评估,结果表明,本方法比当前最先进弱监督基线提升4.3%平均对齐精度,且是唯一超越有监督ROCA基准(+2.7%)的弱监督方法。为评估泛化能力,我们引入SUN2CAD,一个包含20个全新物体类别的真实测试集,本方法在未预先训练的情况下取得最先进性能。
原文摘要 · Abstract (English)
One practical approach to infer 3D scene structure from a single image is to retrieve a closely matching 3D model from a database and align it with the object in the image. Existing methods rely on supervised training with images and pose annotations, which limits them to a narrow set of object categories. To address this, we propose a weakly supervised 9-DoF alignment method for inexact 3D models that requires no pose annotations and generalizes to unseen categories. Our approach derives a novel feature space based on foundation features that ensure multi-view consistency and overcome symmetry ambiguities inherent in foundation features using a self-supervised triplet loss. Additionally, we introduce a texture-invariant pose refinement technique that performs dense alignment in normalized object coordinates, estimated through the enhanced feature space. We conduct extensive evaluations on the real-world ScanNet25k dataset, where our method outperforms SOTA weakly supervised baselines by +4.3% mean alignment accuracy and is the only weakly supervised approach to surpass the supervised ROCA by +2.7%. To assess generalization, we introduce SUN2CAD, a real-world test set with 20 novel object categories, where our method achieves SOTA results without prior training on them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。