arXiv:2607.13110cs.LGcs.AI2026-07中稿 · publication by IEE…

用新型Mamba架构提升音视频导航的动态序列建模能力

A Hybrid Mamba for Audio-Visual Navigation

论文配图:A Hybrid Mamba for Audio-Visual Navigation
图 1 · 摘自论文原文
  • 用自适应选择Mamba替代传统GRU,更好捕捉时序特征
  • 在Matterport3D上导航成功率提升11.3%,复现数据集上更优
  • 适合需要高效多模态感知的机器人导航场景

自2020年以来,音视频导航的基础骨干网络长期依赖卷积神经网络和循环结构,难以有效表征动态多模态序列。本文提出Samba(一种用于音视频导航的混合Mamba模型),采用自适应选择式Mamba状态编码器(M-SE)替代传统GRU进行时间聚合,并构建音频Mamba编码器(AME)以克服卷积操作在频谱图中捕获全局时频依赖性的局限。实验表明,Samba在面对未听过的声源和未知场景时展现出卓越泛化能力:在Matterport3D数据集上,导航成功率(SR)较现有最优模型提升11.3%;在结构更精细的Replica数据集上性能优势更为显著。这种现代化架构重构以更低计算成本实现了更强的具身表征能力,为音视频导航范式演进提供了高鲁棒性技术路径。

原文摘要 · Abstract (English)

Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficient representation of dynamic multimodal sequences. This paper proposes Samba(A Hybrid Mamba for Audio-Visual Navigation). It uses the adaptive selection-enabled Mamba State Encoder (M-SE) to replace conventional GRUs for temporal aggregation, and constructs an Audio Mamba Encoder (AME) to remedy the limitations of convolutional operators in capturing global time-frequency dependencies in spectrograms. Experiments demonstrate that Samba exhibits exceptional generalization performance when facing unheard sound sources and unseen scenes. On the Matterport3D dataset, it improves the navigation success rate (SR) by 11.3\% compared with existing state-of-the-art models, and the performance gain is even more pronounced on the Replica dataset, which features finer scene structures. Such modernized architectural reconstruction unlocks stronger embodied representation capabilities at a lower computational cost, thereby providing a highly robust technical pathway for paradigm evolution in the field of audio-visual navigation.

音视频导航Mamba多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。