用纳米光子金属透镜实现单目图像的物理级深度感知
Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding
- 通过金属透镜在单张图像中编码深度相关的波前偏移
- 在RGB-D数据集上模拟训练,实现毫米级精度的度量深度估计
- 适合需要高精度单目深度感知的机器人与增强现实应用
深度基础模型(DFM)虽能从单张彩色图像中推断3D结构,但缺乏物理尺度信息,导致度量深度不明确。本文提出利用新兴的超薄平面金属透镜,通过纳米光子学物理编码缺失的度量深度线索。在单目成像中,金属透镜将深度依赖的位置偏移嵌入两个偏振光波前。结合输入适配策略,可直接微调预训练的DFM以匹配光学信号。为扩大训练数据,我们构建了综合仿真流程,从RGB-D数据集合成金属透镜响应,引入物理因素以最小化仿真到真实的差距。实验表明,该方法优于单目度量深度估计和基于失焦的深度估计基线,为精确单目度量深度感知提供了有效路径。
原文摘要 · Abstract (English)
Depth foundation models (DFMs) offer strong learned priors for 3D perception from single RGB images but lack physical depth cues, leading to ambiguities in metric scale. We introduce metalenses, an emerging class of ultrathin planar optical elements, as a solution to physically encode missing metric depth cues via nanophotonics. In this paper, we bridge the gap between metalens and DFMs to achieve accurate metric monocular depth sensing. In a single monocular shot, our metalens embeds depth-dependent positional shifts into two polarized optical wavefronts. With an input adaptation strategty, we enable direct fine-tuning that aligns a pretrained DFM with the optical signals. To scale the training data, we further develop a comprehensive simulation pipeline that synthesizes metalens responses from RGB-D datasets, incorporating physical factors to minimize the sim-to-real gap. Experiments demonstrate that this approach outperforms both monocular metric depth estimation and depth-from-defocus baselines, showing an effective pathway for accurate monocular metric depth sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。