arXiv:2602.20157cs.CV2026-02被引 6

用2D光流监督实现无标注视频的高效3D几何学习

Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning

  • 将光流预测分解为几何与位姿两部分,分别从不同图像提取特征
  • 在80万段无标签视频上训练,动态场景表现显著优于现有方法
  • 特别适合真实世界动态视频,标签稀缺时优势更明显

当前前馈式3D/4D重建系统依赖密集的几何与位姿标注,获取成本高且在动态真实场景中尤为稀缺。我们提出Flow3r框架,利用密集2D对应关系(光流)作为监督信号,实现从无标注单目视频中可扩展的视觉几何学习。核心思想是光流预测模块应被分解:使用一张图像的几何隐变量和另一张图像的位姿隐变量进行预测。这种分解直接引导场景几何与相机运动的学习,并自然适用于动态场景。在控制实验中,我们验证了分解式光流预测优于其他设计,且性能随无标签数据量持续提升。将该方法集成到现有视觉几何架构中,使用约80万段无标签视频训练后,Flow3r在涵盖静态与动态场景的八个基准上达到当前最优结果,尤其在真实动态视频上提升最显著,此时标注数据最为稀缺。

原文摘要 · Abstract (English)

Current feed-forward 3D/4D reconstruction systems rely on dense geometry and pose supervision -- expensive to obtain at scale and particularly scarce for dynamic real-world scenes. We present Flow3r, a framework that augments visual geometry learning with dense 2D correspondences (`flow') as supervision, enabling scalable training from unlabeled monocular videos. Our key insight is that the flow prediction module should be factored: predicting flow between two images using geometry latents from one and pose latents from the other. This factorization directly guides the learning of both scene geometry and camera motion, and naturally extends to dynamic scenes. In controlled experiments, we show that factored flow prediction outperforms alternative designs and that performance scales consistently with unlabeled data. Integrating factored flow into existing visual geometry architectures and training with ${\sim}800$K unlabeled videos, Flow3r achieves state-of-the-art results across eight benchmarks spanning static and dynamic scenes, with its largest gains on in-the-wild dynamic videos where labeled data is most scarce.

3D重建光流自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。