用网页视频训练3D估计模型,端到端预测深度、运动和内参。
SS3D: End2End Self-Supervised 3D from Web Videos

- 通过自监督学习联合预测深度、相机运动和内参。
- 在约1亿帧的YouTube-8M数据上预训练,零样本迁移性能强。
- 适合做视觉3D重建或需要少标注数据的场景研究者。
我们提出SS3D,一种基于结构光从单目视频进行3D估计的端到端自监督预训练方法。模型在一次前向传播中联合预测深度、自我运动和相机内参,并作为整体进行训练与评估。为稳定联合学习,采用先优化内参的两阶段训练策略和统一的单检查点评估协议。将结构光自监督扩展至无约束网络视频面临多视图信号弱和语料异质性高的挑战;我们通过多视图信号代理(MVS)实现数据筛选与课程采样,并将专家知识蒸馏至单一学生模型。在过滤后约1亿帧的YouTube-8M数据上预训练后,模型展现出优异的跨域零样本迁移能力,且微调性能优于以往自监督基线。我们已开源预训练模型和代码。
原文摘要 · Abstract (English)
We present SS3D, a web-scale SfM-based self-supervision pretraining pipeline for feed-forward 3D estimation from monocular video. Our model jointly predicts depth, ego-motion, and intrinsics in a single forward pass and is trained/evaluated as a coherent end-to-end 3D estimator. To stabilize joint learning, we use an intrinsics-first two-stage schedule and a unified single-checkpoint evaluation protocol. Scaling SfM self-supervision to unconstrained web video is challenging due to weak multi-view observability and strong corpus heterogeneity; we address these with a multi-view signal proxy (MVS) used for filtering and curriculum sampling, and with expert training distilled into a single student. Pretraining on YouTube-8M (~100M frames after filtering) yields strong cross-domain zero-shot transfer and improved fine-tuning performance over prior self-supervised baselines. We release the pretrained checkpoint and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。