arXiv:2607.02554cs.CVcs.GR2026-07

用可靠区域筛选的单目深度监督,提升稀疏视角神经重建质量

Reliability-Aware Monocular Depth Supervision for Sparse-View Neural Reconstruction

论文配图:Reliability-Aware Monocular Depth Supervision for Sparse-View Neural Reconstruction
图 1 · 摘自论文原文
  • 基于DA-V2生成深度图,通过掩码只在可信区域施加监督
  • Splatfacto模型PSNR提升0.969,RMSE降至0.100,几何更准
  • 适合低重叠多视角场景,尤其对重建精度敏感的应用

室外驾驶场景中稀疏视角神经重建因前向轨迹狭窄、多视角重叠有限而困难。尽管单目深度先验能提供稠密信息,但其可靠性不均且噪声大。本文采用Depth Anything V2(DA-V2)作为稠密单目深度先验,利用稀疏锚点(LiDAR与COLMAP)对每张图像的尺度和偏移进行校准,并通过仅使用RGB的基线模型生成光度掩码,选择性地施加深度监督。在KITTISeq02上,掩码监督对Mip-NeRF-360改善微弱,几何未提升;而Splatfacto显著受益,PSNR从14.903升至15.932,RMSE从0.542降至0.100。相较于全局监督,该掩码在KITTI序列00/02/05上实现0.44-0.70 dB PSNR增益,且保持或优于RMSE。结果表明,掩码主要提升渲染保真度而非几何结构。匹配率消融实验及两个额外KITTI片段验证了收益源于选取低误差可靠区域,而非像素数量减少。在Bicycle场景中,当多视角覆盖已充足时,深度监督虽改善几何却损害渲染质量。以DA-V2为例,结果表明在欠约束稀疏视角重建中,经选择性应用并适度加权的单目深度先验具有价值。

原文摘要 · Abstract (English)

Sparse-view neural reconstruction in outdoor driving is challenging due to narrow forward-facing trajectories and limited multi-view overlap, and monocular depth priors, though dense, are noisy and not uniformly reliable. We use Depth Anything V2 (DA-V2) as a dense monocular depth prior, align its per-image scale and shift to metric depth using sparse anchors (LiDAR and COLMAP) and apply depth supervision selectively through photometric masks generated from an RGB-only baseline model, and evaluate on Mip-NeRF-360 and Splatfacto. On KITTISeq02, masked depth supervision gives only marginal gains for Mip-NeRF-360 and does not improve geometry. In contrast, Splatfacto benefits clearly, improving PSNR from 14.903 to 15.932 and reducing RMSE from 0.542 to 0.100. Against global supervision, the proposed mask achieves 0.44-0.70,dB PSNR gains across KITTI sequences 00/02/05 at tied or better RMSE, while yielding no change on Mip-NeRF-360. This indicates the mask primarily enhances rendering fidelity rather than geometry. Matched-ratio ablations and two further KITTI fragments confirm the gains come from selecting reliable low-error regions, rather than from fewer pixels. On the Bicycle scene, depth supervision improves geometry but hurts RGB rendering quality when multi-view coverage is already strong. Using DA-V2 as a representative prior, results suggest that monocular depth priors are valuable for under-constrained sparse-view reconstruction when applied selectively with moderate weighting.

神经重建单目深度稀疏视图深度监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。