让声音指引视觉,提升机器人听声找物的泛化能力
Audio-Guided Visual Perception for Audio-Visual Navigation
- 用音频上下文引导视觉注意力,实现跨模态对齐
- 在未听过的音源上导航成功率提升37%,路径更短
- 适合需要跨场景适应的智能体导航任务
视听具身导航旨在让智能体在未知三维环境中,仅通过声音线索自主导航至声源位置。现有方法在分布内音源上表现良好,但在遇到未听过的音源或新环境时,导航成功率急剧下降,搜索路径显著变长。其根源在于缺乏声音信号与视觉区域之间的显式对齐机制。策略模型在训练中容易记忆特定的‘声学指纹-场景’关联,导致面对新声音时盲目探索。为此,我们提出AGVP框架,将声音从可记忆的声学指纹转化为空间引导信号。该框架首先通过音频自注意力提取全局音频上下文,再以该上下文作为查询,引导视觉特征注意力,突出与声源相关的视觉区域。随后进行时序建模与策略优化。这一以可解释的跨模态对齐和区域重加权为核心的设计,降低了对特定声学指纹的依赖。实验表明,AGVP在未听过的声音上实现了更高的导航效率与鲁棒性,并显著提升了跨场景泛化能力。
原文摘要 · Abstract (English)
Audio-Visual Embodied Navigation aims to enable agents to autonomously navigate to sound sources in unknown 3D environments using auditory cues. While current AVN methods excel on in-distribution sound sources, they exhibit poor cross-source generalization: navigation success rates plummet and search paths become excessively long when agents encounter unheard sounds or unseen environments. This limitation stems from the lack of explicit alignment mechanisms between auditory signals and corresponding visual regions. Policies tend to memorize spurious \enquote{acoustic fingerprint-scenario} correlations during training, leading to blind exploration when exposed to novel sound sources. To address this, we propose the AGVP framework, which transforms sound from policy-memorable acoustic fingerprint cues into spatial guidance. The framework first extracts global auditory context via audio self-attention, then uses this context as queries to guide visual feature attention, highlighting sound-source-related regions at the feature level. Subsequent temporal modeling and policy optimization are then performed. This design, centered on interpretable cross-modal alignment and region reweighting, reduces dependency on specific acoustic fingerprints. Experimental results demonstrate that AGVP improves both navigation efficiency and robustness while achieving superior cross-scenario generalization on previously unheard sounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。