用视觉模型适配声呐,实现无需姿态信息的3D重建
Adapting Vision Foundation Models to Acoustics for Pose-Free 3D Sonar Reconstruction

- 利用视觉与声呐的几何关系迁移模型
- 通过物理噪声模型生成合成数据提升性能
- 首次实现实测声呐下无姿态3D重建
在互联网规模的RGB数据集上训练的视觉基础模型可实现从文本生成视频到少样本3D场景重建等多种能力。若在大规模声呐数据集上训练声呐基础模型,或可在水下浊度高、能见度低的场景中实现类似能力。但目前缺乏公开的大规模声呐数据集,直接训练不现实。本文提出通过(1)利用两种传感模态间的几何关系,(2)采用精确的物理基噪声模型生成合成数据,高效地将视觉基础模型适配至声呐场景。所提出的适配模型首次在实验中实现了基于声呐的无姿态3D重建,展现出新能力。
原文摘要 · Abstract (English)
Vision foundation models trained on Internet-scale RGB datasets enable remarkable capabilities across a range of tasks, from text-to-video generation to few-shot 3D scene reconstruction. An acoustic foundation model trained on large-scale sonar datasets could enable similar capabilities in the underwater domain, where turbidity and low-visibility conditions make conventional RGB foundation models inapplicable. Unfortunately, a lack of freely available large-scale sonar datasets makes training such a model from scratch impractical. In this work, we demonstrate that vision foundation models can be efficiently adapted to the sonar setting by (1) exploiting the geometric relationship between the two sensing modalities and (2) employing accurate physics-based noise models for synthetic data generation. The resulting sonar adaptation models enable new capabilities: For the first time, we experimentally demonstrate sonar-based pose-free 3D reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。