用单镜头实现百万像素3D成像,突破低纹理、高反光等难题。
Extended monocular 3D imaging
- 融合衍射与偏振信息,通过混合镜头捕捉深度线索。
- 实验实现百万像素级3D点云,无需先验数据且精度高。
- 可识别材质,适用于防伪、目标识别等场景。
三维视觉在机器智能、精密测量等领域至关重要。尽管进展显著,现有三维成像设备仍普遍体积大、复杂度高,且图像分辨率远低于二维设备。许多常见场景中,现有方案常失效。本文提出扩展单目三维成像(EM3D)框架,充分利用光的矢量波特性,通过多阶段融合衍射与偏振深度线索,使用配备衍射-折射混合镜头的紧凑单目相机,实验证明可对传统难处理的远距离场景(如低纹理、高反射或近乎透明)进行毫秒级快照获取,生成百万像素级高精度3D点云,且无需数据先验。此外,发现深度与偏振信息结合可实现独特材质识别能力,为靶标识别、人脸防伪等应用拓展机器智能新可能。该结构简洁强大,为小型化高维机器视觉开辟新路径,推动单目相机在更广泛场景中的部署。
原文摘要 · Abstract (English)
3D vision is of paramount importance for numerous applications ranging from machine intelligence to precision metrology. Despite much recent progress, the majority of 3D imaging hardware remains bulky and complicated and provides much lower image resolution compared to their 2D counterparts. Moreover, there are many well-known scenarios that existing 3D imaging solutions frequently fail. Here, we introduce an extended monocular 3D imaging (EM3D) framework that fully exploits the vectorial wave nature of light. Via the multi-stage fusion of diffraction- and polarization-based depth cues, using a compact monocular camera equipped with a diffractive-refractive hybrid lens, we experimentally demonstrate the snapshot acquisition of a million-pixel and accurate 3D point cloud for extended scenes that are traditionally challenging, including those with low texture, being highly reflective, or nearly transparent, without a data prior. Furthermore, we discover that the combination of depth and polarization information can unlock unique new opportunities in material identification, which may further expand machine intelligence for applications like target recognition and face anti-spoofing. The straightforward yet powerful architecture thus opens up a new path for a higher-dimensional machine vision in a minimal form factor, facilitating the deployment of monocular cameras for applications in much more diverse scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。