利用镜头模糊模式,从手持拍摄视频中恢复深度与稠密轨迹。
Trajectory Densification and Depth from Perspective-based Blur
- 通过分析视频的模糊模式联合估计深度与轨迹
- 在多数据集上实现大范围深度重建,精度优于真实轨迹
- 适合移动端或无稳定器场景下的三维重建应用
在缺乏机械防抖的情况下,手持拍摄时相机不可避免地产生旋转运动,尤其在长曝光条件下引发基于视角的模糊。从光学角度看,这种模糊具有深度-位置依赖性:不同空间位置的物体在相同成像设置下会产生不同的模糊程度。受此启发,我们提出一种新方法,通过分析视频流的模糊模式和稀疏轨迹,联合优化光学算法来估计度量深度。具体而言,使用现成的视觉编码器与点追踪器提取视频信息,通过窗口嵌入与多窗口聚合估计深度图,并利用视觉-语言模型对光学算法生成的稀疏轨迹进行稠密化。在多个深度数据集上的评估表明,该方法在大深度范围内表现强劲,同时具备良好的泛化能力。相较于手持拍摄的真实轨迹,我们的光学算法实现了更高的精度,稠密重建也保持了强准确性。
原文摘要 · Abstract (English)
In the absence of a mechanical stabilizer, the camera undergoes inevitable rotational dynamics during capturing, which induces perspective-based blur especially under long-exposure scenarios. From an optical standpoint, perspective-based blur is depth-position-dependent: objects residing at distinct spatial locations incur different blur levels even under the same imaging settings. Inspired by this, we propose a novel method that estimate metric depth by examining the blur pattern of a video stream and dense trajectory via joint optical design algorithm. Specifically, we employ off-the-shelf vision encoder and point tracker to extract video information. Then, we estimate depth map via windowed embedding and multi-window aggregation, and densify the sparse trajectory from the optical algorithm using a vision-language model. Evaluations on multiple depth datasets demonstrate that our method attains strong performance over large depth range, while maintaining favorable generalization. Relative to the real trajectory in handheld shooting settings, our optical algorithm achieves superior precision and the dense reconstruction maintains strong accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。