用YouTube视频训练3D语义占位,无需人工标注。
YouTube-Occ: Learning Indoor 3D Semantic Occupancy Prediction from YouTube Videos
- 从互联网视频自动构建带语义的3D场景点云
- 在NYUv2和Occ-ScanNet上提升有限数据下的性能
- 适合做自监督3D理解的研究者参考
3D语义占位对精细场景理解至关重要,但在隐私敏感的室内环境中,其发展受限于大规模标注3D数据的缺乏。为此,我们探索从海量、未标定的网络视频中学习室内3D语义占位,同时避免繁琐的人工标注。提出YouTube-Occ,包含自动化数据流水线,利用2D与3D基础模型处理原始网络视频,估计相机几何,重建场景点云,并注入密集语义伪标签。然而,这些合理伪标签在简单监督下无法带来性能提升。为此,我们进一步提出基于特征蒸馏的预训练框架,采用双对齐策略:帧内对齐通过体素锚定高斯化模块,将3D特征与对应2D先验对齐;跨场景对齐通过类别原型蒸馏实现全局语义一致性。实验证明,YouTube-Occ在三个主流架构上均在NYUv2和Occ-ScanNet基准上取得一致提升,尤其在数据受限条件下表现显著。代码与数据将公开,以期推动后续研究。
原文摘要 · Abstract (English)
3D semantic occupancy prediction is crucial for fine-grained scene understanding, yet its advancement in privacy-sensitive indoor environments is fundamentally hindered by the scarcity of large-scale annotated 3D data. To overcome this limitation, we explore learning indoor 3D semantic occupancy prediction from abundant, uncalibrated in-the-wild internet videos while simultaneously bypassing the extensive manual annotation. Specifically, we introduce \textit{YouTube-Occ}, including an automated data pipeline that leverages 2D and 3D foundation models to process raw web videos, estimating camera geometry, reconstructing scene point clouds, and enriching them with dense semantic pseudo-labels. However, these plausible pseudo-labels fail to yield performance gains under naive supervision. To address this impasse, we further propose a pre-training framework driven by feature distillation with a dual-alignment strategy. Within it, an intra-frame alignment utilizes a voxel-anchored Gaussianization module to align 3D features with corresponding 2D priors, whereas a cross-scene alignment achieves global semantic consistency via class-prototype distillation. Empirically, YouTube-Occ delivers consistent gains across three mainstream architectures on the NYUv2 and Occ-ScanNet benchmarks, especially under limited-data conditions. We will publicly release our code and data, hoping to inspire future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。