提出统一光照下三维形状重建新方法,可自动分离光照与法向信息。
Light of Normals: Unified Feature Representation for Universal Photometric Stereo
- 用光照注册令牌和交错注意力块解耦光照与表面法向特征。
- 在真实材质上实现比现有方法更优的三维重建精度与泛化能力。
- 适合需要高精度几何还原的工业检测与数字孪生场景。
通用光度立体视觉(PS)需在任意未知光照条件下运行,且不依赖特定光照模型。当前方法仍面临两大挑战:一是编码器无法保证光照与法向信息解耦;二是高频几何细节易丢失。本文提出LINO UniPS,引入光照注册令牌并配合光照对齐监督,聚合点光源、方向光与环境光;设计交错注意力模块,实现跨图像全局光照联合建模,使编码器能剥离光照影响并保留法向证据。为恢复细粒度几何,采用基于小波的双分支架构与法向梯度感知损失。上述方法构建统一特征空间,其中光照由注册令牌显式表示,法向细节通过小波分支保留。同时构建大规模合成数据集PS-Verse,按几何复杂度与光照多样性分级,并采用从简单到复杂的课程训练策略。大量实验表明,在DiLiGenT、Luces等公开基准上达到新最优性能,对真实材质泛化能力更强,效率更高;消融实验证明光照注册令牌+交错注意力提升特征解耦,小波双分支+法向梯度损失有效恢复细节。
原文摘要 · Abstract (English)
Universal photometric stereo (PS) is defined by two factors: it must (i) operate under arbitrary, unknown lighting conditions and (ii) avoid reliance on specific illumination models. Despite progress (e.g., SDM UniPS), two challenges remain. First, current encoders cannot guarantee that illumination and normal information are decoupled. To enforce decoupling, we introduce LINO UniPS with two key components: (i) Light Register Tokens with light alignment supervision to aggregate point, direction, and environment lights; (ii) Interleaved Attention Block featuring global cross-image attention that takes all lighting conditions together so the encoder can factor out lighting while retaining normal-related evidence. Second, high-frequency geometric details are easily lost. We address this with (i) a Wavelet-based Dual-branch Architecture and (ii) a Normal-gradient Perception Loss. These techniques yield a unified feature space in which lighting is explicitly represented by register tokens, while normal details are preserved via wavelet branch. We further introduce PS-Verse, a large-scale synthetic dataset graded by geometric complexity and lighting diversity, and adopt curriculum training from simple to complex scenes. Extensive experiments show new state-of-the-art results on public benchmarks (e.g., DiLiGenT, Luces), stronger generalization to real materials, and improved efficiency; ablations confirm that Light Register Tokens + Interleaved Attention Block drive better feature decoupling, while Wavelet-based Dual-branch Architecture + Normal-gradient Perception Loss recover finer details.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。