arXiv:2601.08900eess.IVcs.CV2026-01被引 1

首个真实感合成数据集助力单帧光栅投影三维重建,揭示信息缺失是性能瓶颈。

Comprehensive Machine Learning Benchmarking for Fringe Projection Profilometry with Photorealistic Synthetic Data

  • 用NVIDIA Isaac Sim生成15,600张真实感光栅图和300个深度图,构建首个公开基准数据集。
  • 个体归一化使重建精度提升9.1倍,背景光栅对相位参考至关重要不可移除。
  • 混合L1损失与UNet架构表现最佳,但距离传统方法的亚毫米精度仍有差距。

用于光栅投影轮廓测量(FPP)的机器学习方法受限于缺乏大规模、多样化的数据集和标准化评估协议。本文首次提出开源、真实感的合成数据集,基于NVIDIA Isaac Sim生成,包含15,600张光栅图像和300个深度重建结果,覆盖50种物体。研究聚焦单帧FPP,即模型直接从单一光栅图像预测3D深度图,无需时间相位移。通过系统性消融实验,确定长距离(1.5–2.1米)深度预测的最佳学习配置。比较三种深度归一化策略,发现个体归一化(解耦物体形状与绝对尺度)相比原始深度,使重建精度提升9.1倍。进一步表明,移除背景光栅会严重损害所有归一化方法的性能,证明背景光栅提供关键空间相位参考而非噪声。评估六种损失函数,确定混合L1损失为最优。在最佳配置下,对比四种网络架构,发现UNet表现最强,但误差仍远高于经典FPP的亚毫米级精度。各架构间性能差距小,表明主要瓶颈在于信息缺失,而非模型设计:单帧光栅图缺乏足够信息以实现准确深度恢复,除非引入显式相位线索。本工作提供标准化基准,并支持将基于相位的FPP与学习增强相结合的混合方法。数据集发布于https://huggingface.co/datasets/aharoon/fpp-ml-bench,代码见https://github.com/AnushLak/fpp-ml-bench。

原文摘要 · Abstract (English)

Machine learning approaches for fringe projection profilometry (FPP) are hindered by the lack of large, diverse datasets and standardized benchmarking protocols. This paper introduces the first open-source, photorealistic synthetic dataset for FPP, generated using NVIDIA Isaac Sim, comprising 15,600 fringe images and 300 depth reconstructions across 50 objects. We apply this dataset to single-shot FPP, where models predict 3D depth maps directly from individual fringe images without temporal phase shifting. Through systematic ablation studies, we identify optimal learning configurations for long-range (1.5-2.1 m) depth prediction. We compare three depth normalization strategies and show that individual normalization, which decouples object shape from absolute scale, yields a 9.1x improvement in object reconstruction accuracy over raw depth. We further show that removing background fringe patterns severely degrades performance across all normalizations, demonstrating that background fringes provide essential spatial phase reference rather than noise. We evaluate six loss functions and identify Hybrid L1 loss as optimal. Using the best configuration, we benchmark four architectures and find UNet achieves the strongest performance, though errors remain far above the sub-millimeter accuracy of classical FPP. The small performance gap between architectures indicates that the dominant limitation is information deficit rather than model design: single fringe images lack sufficient information for accurate depth recovery without explicit phase cues. This work provides a standardized benchmark and evidence motivating hybrid approaches combining phase-based FPP with learned refinement. The dataset is available at https://huggingface.co/datasets/aharoon/fpp-ml-bench and code at https://github.com/AnushLak/fpp-ml-bench.

三维重建机器学习光栅投影合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。