用相位解码机制打破深度估计中的形状捷径,显著提升精度。
Repairing Shape-Prior Shortcuts in Long-Range Single-Shot Fringe Projection Profilometry

- 输出正弦余弦相位表示,通过固定校准层映射深度,从架构上消除形状捷径
- 深度误差降低至4.46 mm,wrap discontinuity处仅0.103%像素残留误差
- 首次在FPP中实现像素级置信度量化,可定位误差并提升5%以上精度
直接回归深度的单帧条纹投影轮廓术(FPP)网络会利用形状先验捷径,仅从物体边界恢复深度而非依赖条纹相位。在包含15,600张条纹图像、50个物体(距离1.5-2.1米)的逼真合成基准上,最优的UNet基线在14.54毫米对象平均绝对误差(MAE)处饱和,增加数据或模型容量均无法消除该捷径,因假设空间未变。本文提出PhiCalNet,输出包裹相位表示$(\sinϕ, \cosϕ)$并通过固定可微校准层映射至深度,从架构上移除形状先验解法。由于单帧映射非单射,需引入条纹序号作为辅助输入,敏感性分析表明其容忍现实解码误差;与相同物理原理的物理信息神经网络(PINN)基线对比,后者无性能提升,证明架构选择为关键因素。PhiCalNet将对象MAE降低3.3倍至4.46毫米,残差仅占±π跳变处0.103%像素;三帧扩展版本达1.16毫米。两个验证一致:相位是内部最可解释特征,且首次在FPP中实现像素级共形不确定性量化,定位误差于同一跳变处;按快照不一致拒绝前5%像素,使均方根误差下降64%,远超基线的3.5%。
原文摘要 · Abstract (English)
Single-shot fringe projection profilometry (FPP) networks that regress depth directly can exploit a shape-prior shortcut, recovering depth from object boundaries rather than from fringe phase. On a photorealistic synthetic benchmark (15,600 fringe images, 50 objects at 1.5-2.1 m standoff), the best such UNet baseline plateaus at 14.54 mm object mean absolute error (MAE), and neither more data nor more capacity removes the shortcut, because neither changes the hypothesis space the optimizer searches. We introduce PhiCalNet, which outputs a wrapped-phase representation $(\sinϕ, \cosϕ)$ and maps it to depth through a fixed differentiable calibration layer, removing the shape-prior solution architecturally rather than by a loss penalty. Because the single-shot mapping is non-injective without fringe order, PhiCalNet takes the fringe order as auxiliary input, an assumption a sensitivity analysis shows tolerates realistic decoding error; a physics-informed (PINN) baseline with the same physics as a soft penalty yields no gain, isolating the architectural choice as the operative factor. PhiCalNet reduces object MAE 3.3x to 4.46 mm, its residual confined to 0.103% of pixels at the $\pmπ$ wrap discontinuity, and a three-frame extension reaches 1.16 mm. Two checks agree: interpretability makes phase the most decodable internal feature, and pixel-wise conformal uncertainty quantification, to our knowledge the first for FPP, localizes error at the same discontinuity, where rejecting the top 5% of pixels by snapshot disagreement cuts root-mean-square error by 64% versus 3.5% for the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。