arXiv:2412.13848cs.CV2024-12被引 2

融合双摄像头与误差指示,实现手机端高精度深度感知。

MobiFuse: A High-Precision On-device Depth Perception System with Multi-Data Fusion

  • 利用环境物理特性构建深度误差指示模态,量化ToF与立体匹配误差。
  • 通过渐进式融合策略,将几何特征与误差特征结合,深度误差降低77.7%。
  • 适用于3D重建与分割,跨数据集泛化能力强,适合移动端部署。

我们提出MobiFuse,一种在移动设备上实现高精度深度感知的系统,融合双路RGB相机与飞行时间(ToF)相机。为达成此目标,我们基于多种环境因素的物理原理,提出深度误差指示(DEI)模态,用于表征ToF与立体匹配的深度误差。进一步采用渐进式融合策略,将来自ToF和立体深度图的几何特征,与来自DEI模态的深度误差特征融合,生成精确深度图。此外,我们构建了新的ToF-立体深度数据集RealToF,用于模型训练与验证。实验表明,MobiFuse相比基线方法显著降低深度测量误差,最高达77.7%。其在多个不同数据集上均展现强泛化能力,并在3D重建与3D分割两个下游任务中证明有效性。MobiFuse在真实场景中的演示视频可通过去标识化的YouTube链接查看:https://youtu.be/jy-Sp7T1LVs。

原文摘要 · Abstract (English)

We present MobiFuse, a high-precision depth perception system on mobile devices that combines dual RGB and Time-of-Flight (ToF) cameras. To achieve this, we leverage physical principles from various environmental factors to propose the Depth Error Indication (DEI) modality, characterizing the depth error of ToF and stereo-matching. Furthermore, we employ a progressive fusion strategy, merging geometric features from ToF and stereo depth maps with depth error features from the DEI modality to create precise depth maps. Additionally, we create a new ToF-Stereo depth dataset, RealToF, to train and validate our model. Our experiments demonstrate that MobiFuse excels over baselines by significantly reducing depth measurement errors by up to 77.7%. It also showcases strong generalization across diverse datasets and proves effectiveness in two downstream tasks: 3D reconstruction and 3D segmentation. The demo video of MobiFuse in real-life scenarios is available at the de-identified YouTube link(https://youtu.be/jy-Sp7T1LVs).

深度感知多模态融合移动端3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。