无需相机参数,实现视频立体匹配的高精度与时间一致性。
Stereo Any Video: Temporally Consistent Stereo Matching
- 利用单目视频深度模型先验,融合卷积特征生成稳定表示。
- 通过全对全相关和时序凸上采样,提升匹配精度与时间连贯性。
- 零样本下跨数据集表现优异,适用于真实室内外场景。
本文提出 Stereo Any Video,一种强大的视频立体匹配框架,无需依赖相机位姿或光流等辅助信息,即可实现空间精确且时间一致的视差估计。其强大性能源于单目视频深度模型提供的丰富先验,并与卷积特征结合,生成稳定表征。关键设计包括:全对全对相关(all-to-all-pairs correlation),构建平滑鲁棒的匹配代价体;以及时序凸上采样(temporal convex upsampling),增强时间一致性。大量实验表明,该方法在多个数据集上均达到领先水平,且在零样本设置下具备强泛化能力,可有效适应真实室内外环境。
原文摘要 · Abstract (English)
This paper introduces Stereo Any Video, a powerful framework for video stereo matching. It can estimate spatially accurate and temporally consistent disparities without relying on auxiliary information such as camera poses or optical flow. The strong capability is driven by rich priors from monocular video depth models, which are integrated with convolutional features to produce stable representations. To further enhance performance, key architectural innovations are introduced: all-to-all-pairs correlation, which constructs smooth and robust matching cost volumes, and temporal convex upsampling, which improves temporal coherence. These components collectively ensure robustness, accuracy, and temporal consistency, setting a new standard in video stereo matching. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple datasets both qualitatively and quantitatively in zero-shot settings, as well as strong generalization to real-world indoor and outdoor scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。