arXiv:2508.13091cs.CV2025-08ICCV被引 3

用扩散模型生成新视图,提升自监督深度估计精度。

DMS:Diffusion-Based Multi-Baseline Stereo Generation for Improving Self-Supervised Depth Estimation

  • 用扩散模型生成左-左、右-右及中间新视角,补全遮挡区域。
  • 在多个数据集上实现最高性能,异常点减少35%。
  • 无需标签数据,可直接接入现有自监督方法使用。

尽管基于学习的立体匹配和单目深度估计已取得显著进展,但利用立体图像作为监督信号的自监督方法仍受关注较少,且需进一步研究。主要挑战来自光度重建中的模糊性,尤其在目标视角的遮挡和画面外区域缺失对应像素时。为此,我们提出DMS,一种模型无关的方法,利用扩散模型的几何先验,在视差方向合成新视角,由方向提示引导。具体而言,微调Stable Diffusion模型以模拟关键位置视角:从左相机偏移的左-左视图、从右相机偏移的右-右视图,以及左右相机间的额外新视图。这些合成视图补充了遮挡像素,实现明确的光度重建。所提DMS为零成本、即插即用方法,能无缝增强自监督立体匹配与单目深度估计,仅依赖未标注的立体图像对进行训练与合成。大量实验表明其有效性,异常点最多减少35%,并在多个基准数据集上达到领先性能。

原文摘要 · Abstract (English)

While supervised stereo matching and monocular depth estimation have advanced significantly with learning-based algorithms, self-supervised methods using stereo images as supervision signals have received relatively less focus and require further investigation. A primary challenge arises from ambiguity introduced during photometric reconstruction, particularly due to missing corresponding pixels in ill-posed regions of the target view, such as occlusions and out-of-frame areas. To address this and establish explicit photometric correspondences, we propose DMS, a model-agnostic approach that utilizes geometric priors from diffusion models to synthesize novel views along the epipolar direction, guided by directional prompts. Specifically, we finetune a Stable Diffusion model to simulate perspectives at key positions: left-left view shifted from the left camera, right-right view shifted from the right camera, along with an additional novel view between the left and right cameras. These synthesized views supplement occluded pixels, enabling explicit photometric reconstruction. Our proposed DMS is a cost-free, ''plug-and-play'' method that seamlessly enhances self-supervised stereo matching and monocular depth estimation, and relies solely on unlabeled stereo image pairs for both training and synthesizing. Extensive experiments demonstrate the effectiveness of our approach, with up to 35% outlier reduction and state-of-the-art performance across multiple benchmark datasets.

自监督深度估计扩散模型立体匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。