arXiv:2512.14028cs.CV2025-12SIGGRAPH被引 3

用神经特征解码提升单帧结构光3D成像的鲁棒性

Robust Single-shot Structured Light 3D Imaging via Neural Feature Decoding

  • 在特征空间而非像素域进行匹配,增强对遮挡等挑战的适应性
  • 合成数据训练下仍可泛化到真实场景,精度超越商业系统
  • 适合需要高鲁棒性3D感知的工业与消费级应用

针对单帧结构光系统在遮挡、细结构和非朗伯表面等挑战下的性能瓶颈,本文提出一种基于神经特征解码的3D成像方法。通过提取投影图案与红外图像的神经特征,并在特征空间构建代价体以引入几何先验,实现更鲁棒的对应关系匹配。进一步设计深度优化模块,利用大规模单目深度估计模型的强先验,提升细节恢复与整体结构一致性。为支持有效学习,构建物理驱动的结构光渲染流水线,生成近百万对合成图像-模式对,覆盖多种物体与材质。实验表明,该方法仅在合成数据上训练,即可在真实室内环境中泛化良好,支持多种模式无需重训,且持续优于商用结构光系统及基于被动立体视觉的深度估计方法。

原文摘要 · Abstract (English)

We consider the problem of active 3D imaging using single-shot structured light systems, which are widely employed in commercial 3D sensing devices such as Apple Face ID and Intel RealSense. Traditional structured light methods typically decode depth correspondences through pixel-domain matching algorithms, resulting in limited robustness under challenging scenarios like occlusions, fine-structured details, and non-Lambertian surfaces. Inspired by recent advances in neural feature matching, we propose a learning-based structured light decoding framework that performs robust correspondence matching within feature space rather than the fragile pixel domain. Our method extracts neural features from the projected patterns and captured infrared (IR) images, explicitly incorporating their geometric priors by building cost volumes in feature space, achieving substantial performance improvements over pixel-domain decoding approaches. To further enhance depth quality, we introduce a depth refinement module that leverages strong priors from large-scale monocular depth estimation models, improving fine detail recovery and global structural coherence. To facilitate effective learning, we develop a physically-based structured light rendering pipeline, generating nearly one million synthetic pattern-image pairs with diverse objects and materials for indoor settings. Experiments demonstrate that our method, trained exclusively on synthetic data with multiple structured light patterns, generalizes well to real-world indoor environments, effectively processes various pattern types without retraining, and consistently outperforms both commercial structured light systems and passive stereo RGB-based depth estimation methods. Project page: https://namisntimpot.github.io/NSLweb/.

3D成像结构光神经特征深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。