arXiv:2509.21377cs.CVcs.AI2025-09中稿 · publication by ECA…被引 7

动态融合多目标视听信息,让机器人更准更快找到声源。

Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation

  • 用改进的Transformer动态筛选并融合视觉与听觉信息
  • 在Replica和Matterport3D上成功率、路径效率均达新高
  • 适合研究多模态机器人导航的开发者与工程师

视听具身导航使机器人通过集成机载传感器的视觉信息与目标发出的音频信号,动态定位声源。核心挑战在于有效利用多模态线索引导导航。尽管已有工作探索了视觉与音频数据的基本融合,但常忽略深层感知上下文。为此,我们提出高效视听导航的动态多目标融合方法(DMTF-AVN)。该方法采用多目标架构与优化的Transformer机制,实现跨模态信息的过滤与选择性融合。在Replica和Matterport3D数据集上的大量实验表明,DMTF-AVN在成功率(SR)、路径效率(SPL)和场景适应性(SNA)方面均达到当前最优表现。此外,模型展现出强可扩展性与泛化能力,为机器人导航中的高级多模态融合策略提供新路径。代码与视频见https://github.com/zzzmmm-svg/DMTF。

原文摘要 · Abstract (English)

Audiovisual embodied navigation enables robots to locate audio sources by dynamically integrating visual observations from onboard sensors with the auditory signals emitted by the target. The core challenge lies in effectively leveraging multimodal cues to guide navigation. While prior works have explored basic fusion of visual and audio data, they often overlook deeper perceptual context. To address this, we propose the Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation (DMTF-AVN). Our approach uses a multi-target architecture coupled with a refined Transformer mechanism to filter and selectively fuse cross-modal information. Extensive experiments on the Replica and Matterport3D datasets demonstrate that DMTF-AVN achieves state-of-the-art performance, outperforming existing methods in success rate (SR), path efficiency (SPL), and scene adaptation (SNA). Furthermore, the model exhibits strong scalability and generalizability, paving the way for advanced multimodal fusion strategies in robotic navigation. The code and videos are available at https://github.com/zzzmmm-svg/DMTF.

视听导航多模态融合机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。