发现远距离3D扫描中模型依赖形状先验而非相位信息,导致误差随距离平方增长。
Diagnosing Shape-Prior Shortcuts in Long-Range Single-Shot Fringe Projection Profilometry

- 通过可解释性分析定位模型依赖物体边界而非相位解码
- 长距离下错误条纹序导致深度误差与距离平方成正比
- 现有数据和模型规模无法消除形状捷径,需重构网络结构
基于学习的单帧条纹投影轮廓术(FPP)研究几乎仅限于近距离场景,且网络性能仅以整体误差评估,未明确其深度恢复是基于条纹相位还是与深度相关的物体形状线索。本文在远距离(超过1米)条件下机制性地诊断该问题。利用公开的、高保真合成基准FPP-ML-Bench(包含15,600张条纹图像,50个物体位于1.5–2.1米距离),首先形式化说明:在远距离下,单帧条纹到深度的映射更严重病态——缺乏条纹序信息时非单射,且错误条纹序引起的深度误差随工作距离的平方 $Z^2$ 增长。系统性消融实验结合多帧研究建立最优UNet基线,达到14.54毫米的物体平均绝对误差(MAE),占80毫米物体深度范围的18%,四种架构间仅1.9倍差异,表明为表示能力限制而非容量瓶颈。首个应用于FPP网络的机制可解释性研究显示:线性探测表明边缘可解码度是深度的2.82倍;Grad-CAM显示注意力对边界的偏好比条纹高出1.28倍;平面测试中,即使条纹有效,无特征平面仍被误判为背景深度。基线模型通过物体边界形状先验完成任务,而非条纹相位解码。由于该捷径是假设空间属性,增加数据或更大模型无法消除,因而提出通过结构设计直接移除形状先验解法。
原文摘要 · Abstract (English)
Learning-based single-shot fringe projection profilometry (FPP) has been studied almost entirely at close range, and the networks used are evaluated only on aggregate error, leaving open whether they recover depth from fringe phase or from object-level shape cues that correlate with depth. This paper diagnoses that question mechanistically in the long-range regime (standoff beyond 1 m). Using FPP-ML-Bench, an open photorealistic synthetic benchmark (15,600 fringe images, 50 objects at 1.5--2.1 m), we first formalize why the single-shot fringe-to-depth mapping is more severely ill-posed at long range: it is non-injective without fringe-order information, and the depth error from an incorrect fringe order grows as $Z^2$ in the working distance. Systematic ablations, extended with a multi-frame study, establish a best UNet baseline at 14.54 mm object mean absolute error (MAE), 18% of the 80 mm object depth range, with only a 1.9$\times$ spread across four architectures, indicating a representational rather than a capacity-bound limit. A mechanistic interpretability study, the first applied to an FPP network, localizes the cause: linear probing shows edges are 2.82$\times$ more decodable than depth, Grad-CAM shows attention favoring boundaries over fringes by 1.28$\times$, and an in-range flat-plane test collapses a featureless plane to background depth despite valid fringes. The baseline solves the task via object-boundary shape priors rather than fringe-phase decoding. Because the shortcut is a hypothesis-space property, additional data or larger models will not remove it, motivating an architectural repair that removes the shape-prior solution by construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。