无需标注数据,用自监督方法完成单图3D场景语义重建
Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
- 基于自监督学习的前馈式3D特征生成,无需地面真值
- 在无监督场景理解中达到顶尖分割精度,3D特征线性探测媲美有监督模型
- 适合关注无监督3D理解、跨域泛化能力的研究者
语义场景补全(SSC)旨在从单张图像中推断场景的三维几何与语义信息。与以往依赖昂贵标注的方法不同,本文提出无监督设置下的新方法SceneDINO,借鉴自监督表示学习和2D无监督场景理解技术。训练过程仅使用多视角一致性自监督信号,不依赖任何语义或几何真值。给定单张输入图像,SceneDINO以前馈方式推断3D几何结构与丰富的3D DINO特征。通过创新的3D特征蒸馏策略,实现无监督3D语义获取。在3D与2D无监督场景理解任务中,SceneDINO均达到当前最优分割准确率;其3D特征经线性探测后,分割性能可比肩现有有监督的SSC方法。此外,本文展示了SceneDINO在跨域泛化与多视角一致性方面的优势,为单图3D场景理解奠定坚实基础。
原文摘要 · Abstract (English)
Semantic scene completion (SSC) aims to infer both the 3D geometry and semantics of a scene from single images. In contrast to prior work on SSC that heavily relies on expensive ground-truth annotations, we approach SSC in an unsupervised setting. Our novel method, SceneDINO, adapts techniques from self-supervised representation learning and 2D unsupervised scene understanding to SSC. Our training exclusively utilizes multi-view consistency self-supervision without any form of semantic or geometric ground truth. Given a single input image, SceneDINO infers the 3D geometry and expressive 3D DINO features in a feed-forward manner. Through a novel 3D feature distillation approach, we obtain unsupervised 3D semantics. In both 3D and 2D unsupervised scene understanding, SceneDINO reaches state-of-the-art segmentation accuracy. Linear probing our 3D features matches the segmentation accuracy of a current supervised SSC approach. Additionally, we showcase the domain generalization and multi-view consistency of SceneDINO, taking the first steps towards a strong foundation for single image 3D scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。