用互联网视频训练3D模型,让几何基础模型更通用。
Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video
- 从视频中提取关键轨迹,结合稀疏与密集监督信号。
- 在多个未见数据集上,3D重建误差降低20%-42%。
- 适合想低成本提升3D模型泛化能力的研究者。
几何基础模型在3D重建中展现出潜力,但其发展受限于多样且大规模的3D标注数据稀缺。尽管互联网视频提供近乎无限的原始数据,但由于缺乏真实几何信息和观测噪声的存在,难以直接用于几何学习。为此,我们提出SAGE框架,实现从原始视频流中可扩展地适应几何基础模型。SAGE采用分层挖掘流程,将视频转化为训练轨迹,并引入混合监督:(1) 信息性训练轨迹选择;(2) 基于SfM点云的稀疏几何锚定,提供全局结构引导;(3) 基于3D高斯渲染的稠密可微一致性,实现多视角约束。为防止灾难性遗忘,引入锚点数据正则化策略。大量实验表明,SAGE显著提升零样本泛化性能,在未见基准(7Scenes、TUM-RGBD、Matterport3D)上,切比雪夫距离降低20%-42%,优于当前最优基线。据我们所知,SAGE首次通过互联网视频实现几何基础模型的适应,建立了一种通用3D学习的可扩展范式。
原文摘要 · Abstract (English)
Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet videos offer virtually unlimited raw data, utilizing them as a scaling source for geometric learning is challenging due to the absence of ground-truth geometry and the presence of observational noise. To address this, we propose SAGE, a framework for Scalable Adaptation of GEometric foundation models from raw video streams. SAGE leverages a hierarchical mining pipeline to transform videos into training trajectories and hybrid supervision: (1) Informative training trajectory selection; (2) Sparse Geometric Anchoring via SfM point clouds for global structural guidance; and (3) Dense Differentiable Consistency via 3D Gaussian rendering for multi-view constraints. To prevent catastrophic forgetting, we introduce a regularization strategy using anchor data. Extensive experiments show that SAGE significantly enhances zero-shot generalization, reducing Chamfer Distance by 20-42% on unseen benchmarks (7Scenes, TUM-RGBD, Matterport3D) compared to state-of-the-art baselines. To our knowledge, SAGE pioneers the adaptation of geometric foundation models via Internet video, establishing a scalable paradigm for general-purpose 3D learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。