arXiv:2509.16924cs.AIcs.SD2025-09中稿 · publication by ICO…被引 10

通过立体声空间感知与动态融合提升声音导航精度

Audio-Guided Dynamic Modality Fusion with Stereo-Aware Attention for Audio-Visual Navigation

  • 利用左右声道差异建模声音方向,增强立体感知
  • 音频驱动动态融合视觉与听觉特征,提升环境适应性
  • 在复杂场景下表现优异,适合机器人导航应用

在音频-视觉导航(AVN)任务中,具身智能体需基于音视频信号,在未知复杂的3D环境中自主定位声源。现有方法多采用静态模态融合策略,忽略立体音频中的空间线索,导致在杂乱或遮挡场景下性能下降。为此,我们提出一种基于强化学习的端到端AVN框架,包含两项关键创新:(1) 立体感知注意力模块(SAM),通过学习左右声道间的空间差异,增强方向性声音感知;(2) 音频引导动态融合模块(AGDF),根据音频线索动态调整视觉与听觉特征的融合比例,提升对环境变化的鲁棒性。在两个真实3D场景数据集Replica和Matterport3D上进行大量实验,结果表明,本方法在导航成功率和路径效率上均显著优于现有方法。尤其在仅音频条件下,相比最佳基线模型,成功率达40%以上提升。结果凸显了显式建模立体声空间线索与深度多模态融合对实现鲁棒高效音频-视觉导航的重要性。

原文摘要 · Abstract (English)

In audio-visual navigation (AVN) tasks, an embodied agent must autonomously localize a sound source in unknown and complex 3D environments based on audio-visual signals. Existing methods often rely on static modality fusion strategies and neglect the spatial cues embedded in stereo audio, leading to performance degradation in cluttered or occluded scenes. To address these issues, we propose an end-to-end reinforcement learning-based AVN framework with two key innovations: (1) a \textbf{S}tereo-Aware \textbf{A}ttention \textbf{M}odule (\textbf{SAM}), which learns and exploits the spatial disparity between left and right audio channels to enhance directional sound perception; and (2) an \textbf{A}udio-\textbf{G}uided \textbf{D}ynamic \textbf{F}usion Module (\textbf{AGDF}), which dynamically adjusts the fusion ratio between visual and auditory features based on audio cues, thereby improving robustness to environmental changes. Extensive experiments are conducted on two realistic 3D scene datasets, Replica and Matterport3D, demonstrating that our method significantly outperforms existing approaches in terms of navigation success rate and path efficiency. Notably, our model achieves over 40\% improvement under audio-only conditions compared to the best-performing baselines. These results highlight the importance of explicitly modeling spatial cues from stereo channels and performing deep multi-modal fusion for robust and efficient audio-visual navigation.

音频导航多模态融合立体声感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。