用声音位置信息动态融合视听数据,提升导航泛化能力
Audio Spatially-Guided Fusion for Audio-Visual Navigation
- 通过声音强度注意力机制提取空间特征
- 在Replica和Matterport3D上对未知声源任务提升泛化性能
- 适合研究多模态机器人导航的开发者
视听导航指智能体在复杂三维环境中利用视觉与听觉信息完成目标定位与路径规划,实现自主导航。该任务的核心挑战在于:如何使智能体摆脱对训练数据的依赖,在环境与声源变化时仍具备良好泛化能力。为此,我们提出一种音频空间引导的视听融合方法。首先设计音频空间特征编码器,通过音频强度注意力机制自适应提取与目标相关联的空间状态信息;基于此,引入音频空间状态引导融合(ASGF)模块,实现多模态特征的动态对齐与自适应融合,有效缓解感知不确定性带来的噪声干扰。在Replica和Matterport3D数据集上的实验表明,该方法在未见过的任务中表现优异,尤其在未知声源分布下展现出更强的泛化能力。
原文摘要 · Abstract (English)
Audio-visual Navigation refers to an agent utilizing visual and auditory information in complex 3D environments to accomplish target localization and path planning, thereby achieving autonomous navigation. The core challenge of this task lies in the following: how the agent can break free from the dependence on training data and achieve autonomous navigation with good generalization performance when facing changes in environments and sound sources. To address this challenge, we propose an Audio Spatially-Guided Fusion for Audio-Visual Navigation method. First, we design an audio spatial feature encoder, which adaptively extracts target-related spatial state information through an audio intensity attention mechanism; based on this, we introduce an Audio Spatial State Guided Fusion (ASGF) to achieve dynamic alignment and adaptive fusion of multimodal features, effectively alleviating noise interference caused by perceptual uncertainty. Experimental results on the Replica and Matterport3D datasets indicate that our method is particularly effective on unheard tasks, demonstrating improved generalization under unknown sound source distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。