arXiv:2411.17249cs.CVcs.AI2024-11CVPR被引 13

无需配对数据,用单图先验实现视频深度与法向量的高质量估计

Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors

  • 基于单图先验和时序一致性约束,构建零样本视频几何估计框架
  • 在无配对视频数据条件下,时序一致性显著提升且精度接近顶尖视频模型
  • 适用于缺乏标注视频数据的场景,适合快速部署于现有图像模型

我们提出 Buffer Anytime 框架,实现从视频中估计深度图与法向量图(统称几何缓冲),无需依赖成对的视频-深度或视频-法向量训练数据。不依赖大规模标注视频数据集,而是通过单图先验结合时序一致性约束,实现高质量视频缓冲估计。采用基于光流平滑性的先进图像估计模型,通过混合损失函数与轻量级时序注意力架构实现零样本训练策略。应用于 Depth Anything V2 和 Marigold-E2E-FT 等领先图像模型时,显著提升时序一致性,同时保持高精度。实验表明,该方法不仅优于纯图像基方法,且在未使用任何成对视频数据的前提下,性能可媲美基于大规模配对视频数据训练的最先进视频模型。

原文摘要 · Abstract (English)

We present Buffer Anytime, a framework for estimation of depth and normal maps (which we call geometric buffers) from video that eliminates the need for paired video--depth and video--normal training data. Instead of relying on large-scale annotated video datasets, we demonstrate high-quality video buffer estimation by leveraging single-image priors with temporal consistency constraints. Our zero-shot training strategy combines state-of-the-art image estimation models based on optical flow smoothness through a hybrid loss function, implemented via a lightweight temporal attention architecture. Applied to leading image models like Depth Anything V2 and Marigold-E2E-FT, our approach significantly improves temporal consistency while maintaining accuracy. Experiments show that our method not only outperforms image-based approaches but also achieves results comparable to state-of-the-art video models trained on large-scale paired video datasets, despite using no such paired video data.

视频深度几何估计零样本时序一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。