让单目视频深度图时间一致,无需训练视频模型
Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
- 用DUSt3R对齐不同时刻的单目深度图,实现时序一致性
- 可同时恢复深度图与相机位姿,性能优于基线方法
- 无需训练视频扩散模型,适合动态场景深度估计
近期单目深度估计方法虽能高质量还原单帧图像深度,但在视频序列中难以保持帧间一致性。现有方法多依赖视频扩散模型生成条件深度,但训练成本高且仅输出尺度不变的深度值,无法获取相机位姿。本文提出新方法Align3R,通过微调DUSt3R模型,将估计的单目深度作为输入,对动态场景下的不同时间步深度图进行对齐。首先在动态场景上微调DUSt3R,再通过优化联合重建深度图与相机位姿。大量实验表明,Align3R在单目视频上实现了更优的时间一致性深度估计与相机位姿恢复,优于基线方法。
原文摘要 · Abstract (English)
Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video diffusion model to generate video depth conditioned on the input video, which is training-expensive and can only produce scale-invariant depth values without camera poses. In this paper, we propose a novel video-depth estimation method called Align3R to estimate temporal consistent depth maps for a dynamic video. Our key idea is to utilize the recent DUSt3R model to align estimated monocular depth maps of different timesteps. First, we fine-tune the DUSt3R model with additional estimated monocular depth as inputs for the dynamic scenes. Then, we apply optimization to reconstruct both depth maps and camera poses. Extensive experiments demonstrate that Align3R estimates consistent video depth and camera poses for a monocular video with superior performance than baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。