arXiv:2512.23042cs.CV2025-12中稿 · CVPR被引 1

用视频生成点云实现大规模3D自监督学习,无需真实扫描数据

3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds

  • 基于视频重建点云,构建无真实扫描的自监督训练框架
  • 在49,219个视频生成场景上训练,室内语义分割性能超越以往方法
  • 适合研究3D自监督学习与数据高效建模的学者使用

尽管3D自监督学习取得进展,但大规模3D场景扫描仍成本高昂。本文探索是否可仅从无真实3D传感器的未标注视频中学习3D表示。提出LAM3C框架,利用视频生成点云进行自监督学习。构建了RoomTours数据集,从网络收集实景看房视频(如房产展示),通过现成前馈重建模型生成49,219个场景点云。设计噪声正则化损失,强化局部几何平滑性,提升点云噪声下的特征稳定性。令人惊讶的是,不依赖任何真实3D扫描,该方法在室内语义与实例分割任务上优于以往自监督模型。结果表明,未标注视频是3D自监督学习的丰富数据来源。代码已开源。

原文摘要 · Abstract (English)

Despite recent progress in 3D self-supervised learning, collecting large-scale 3D scene scans remains expensive and labor-intensive. In this work, we investigate whether 3D representations can be learned from unlabeled videos recorded without any real 3D sensors. We present Laplacian-Aware Multi-level 3D Clustering with Sinkhorn-Knopp (LAM3C), a self-supervised framework that learns from video-generated point clouds reconstructed from unlabeled videos. We first introduce RoomTours, a video-generated point cloud dataset constructed by collecting room-walkthrough videos from the web (e.g., real-estate tours) and generating 49,219 scenes using an off-the-shelf feed-forward reconstruction model. We also propose a noise-regularized loss that stabilizes representation learning by enforcing local geometric smoothness and ensuring feature stability under noisy point clouds. Remarkably, without using any real 3D scans, LAM3C achieves better performance than previous self-supervised methods on indoor semantic and instance segmentation. These results suggest that unlabeled videos represent an abundant source of data for 3D self-supervised learning. Our source code is available at https://ryosuke-yamada.github.io/lam3c/.

3D自监督视频生成点云重建数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。