用声音和视觉融合重建室内透明表面,低成本高效精准。
VAIR: Visuo-Acoustic Implicit Representations for Low-Cost, Multi-Modal Transparent Surface Reconstruction in Indoor Scenes
- 通过隐式神经表示融合声学与视觉信息
- 在新采集的数据集上显著优于现有方法
- 适合移动机器人在复杂室内环境导航
移动机器人在室内环境中需应对包含透明表面的复杂场景。本文提出一种新方法,通过隐式神经表示融合声学与视觉传感模态,实现室内场景中透明表面的稠密重建。我们设计了一种基于生成潜在优化的新模型,学习包含透明表面的室内场景隐式表示。可通过查询该表示,在图像空间实现体渲染,或重建点云与网格,并预测透明表面。我们在自研的低成本感知平台(配备RGB-D相机与超声传感器)采集的新数据集上进行评估,结果表明该方法在定性与定量层面均显著优于当前最优水平。
原文摘要 · Abstract (English)
Mobile robots operating indoors must be prepared to navigate challenging scenes that contain transparent surfaces. This paper proposes a novel method for the fusion of acoustic and visual sensing modalities through implicit neural representations to enable dense reconstruction of transparent surfaces in indoor scenes. We propose a novel model that leverages generative latent optimization to learn an implicit representation of indoor scenes consisting of transparent surfaces. We demonstrate that we can query the implicit representation to enable volumetric rendering in image space or 3D geometry reconstruction (point clouds or mesh) with transparent surface prediction. We evaluate our method's effectiveness qualitatively and quantitatively on a new dataset collected using a custom, low-cost sensing platform featuring RGB-D cameras and ultrasonic sensors. Our method exhibits significant improvement over state-of-the-art for transparent surface reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。