利用几何先验提升复杂场景下的光照恢复能力
Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo under Limited Multi-Illumination Cues
- 融合3D重建模型的几何先验与多光照线索
- 在真实透视投影下实现更准确的表面法向估计
- 特别适合复杂野外场景,性能优于现有方法
通用光度立体法是一种无需严格光照假设即可恢复表面法向的有前景方法。然而,在光照偏置、阴影或复杂野外场景的自遮挡区域中,多光照线索不可靠时,该方法表现不佳。本文提出GeoUniPS,一种集成合成监督与大规模3D重建模型所预训练的高层几何先验的通用光度立体网络。核心思想是:这些3D重建模型作为视觉-几何基础模型,天然蕴含真实场景的丰富几何知识。为此,设计了光照-几何双分支编码器,从冻结的3D重建模型中提取多光照线索和几何先验。同时,为克服传统正交投影假设的局限性,引入包含真实透视投影的PS-Perp数据集,以学习空间变化的视角方向。大量实验表明,GeoUniPS在多个数据集上均达到当前最优性能,尤其在复杂野外场景中表现突出。
原文摘要 · Abstract (English)
Universal Photometric Stereo is a promising approach for recovering surface normals without strict lighting assumptions. However, it struggles when multi-illumination cues are unreliable, such as under biased lighting or in shadows or self-occluded regions of complex in-the-wild scenes. We propose GeoUniPS, a universal photometric stereo network that integrates synthetic supervision with high-level geometric priors from large-scale 3D reconstruction models pretrained on massive in-the-wild data. Our key insight is that these 3D reconstruction models serve as visual-geometry foundation models, inherently encoding rich geometric knowledge of real scenes. To leverage this, we design a Light-Geometry Dual-Branch Encoder that extracts both multi-illumination cues and geometric priors from the frozen 3D reconstruction model. We also address the limitations of the conventional orthographic projection assumption by introducing the PS-Perp dataset with realistic perspective projection to enable learning of spatially varying view directions. Extensive experiments demonstrate that GeoUniPS delivers state-of-the-arts performance across multiple datasets, both quantitatively and qualitatively, especially in the complex in-the-wild scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。