用海量无姿态视频训练3D生成模型,看一眼就能学会造3D。
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale

- 从互联网视频中自动筛选多视角一致数据,构建320M帧的大规模多视图数据集
- 通过添加时间噪声生成纯2D视觉信号,无需相机姿态也能学习3D结构
- 零样本、开放世界生成能力强,性能超越依赖昂贵3D标注的现有模型
近期3D生成模型通常依赖有限规模的3D标注数据或2D扩散先验,其性能受限于受限的3D先验。本文提出See3D,一种基于大规模网络视频的视觉条件多视角扩散模型,实现开放世界的3D内容生成。模型仅通过观察海量快速增长的视频内容即可获取3D知识——‘你看到它,你就拥有它’。我们设计了一套数据清洗流程,自动剔除多视角不一致与观测不足的视频片段,构建高质量、多样化且大规模的多视图图像数据集WebVi3D,包含1600万视频片段中的3.2亿帧。尽管缺乏显式3D几何或相机姿态标注,学习通用3D先验极具挑战,而对网络视频进行姿态标注成本过高。为此,我们引入一种创新的纯2D诱导视觉信号:对掩码视频数据添加时间相关噪声。最终,我们将See3D集成到基于变形的流水线中,实现高保真3D生成。在单视图和稀疏重建基准上的数值与视觉对比显示,仅使用低成本、可扩展的视频数据训练的See3D,在零样本与开放世界生成任务中表现显著优于依赖昂贵且受限3D数据集的模型。
原文摘要 · Abstract (English)
Recent 3D generation models typically rely on limited-scale 3D `gold-labels' or 2D diffusion priors for 3D content creation. However, their performance is upper-bounded by constrained 3D priors due to the lack of scalable learning paradigms. In this work, we present See3D, a visual-conditional multi-view diffusion model trained on large-scale Internet videos for open-world 3D creation. The model aims to Get 3D knowledge by solely Seeing the visual contents from the vast and rapidly growing video data -- You See it, You Got it. To achieve this, we first scale up the training data using a proposed data curation pipeline that automatically filters out multi-view inconsistencies and insufficient observations from source videos. This results in a high-quality, richly diverse, large-scale dataset of multi-view images, termed WebVi3D, containing 320M frames from 16M video clips. Nevertheless, learning generic 3D priors from videos without explicit 3D geometry or camera pose annotations is nontrivial, and annotating poses for web-scale videos is prohibitively expensive. To eliminate the need for pose conditions, we introduce an innovative visual-condition - a purely 2D-inductive visual signal generated by adding time-dependent noise to the masked video data. Finally, we introduce a novel visual-conditional 3D generation framework by integrating See3D into a warping-based pipeline for high-fidelity 3D generation. Our numerical and visual comparisons on single and sparse reconstruction benchmarks show that See3D, trained on cost-effective and scalable video data, achieves notable zero-shot and open-world generation capabilities, markedly outperforming models trained on costly and constrained 3D datasets. Please refer to our project page at: https://vision.baai.ac.cn/see3d
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。