arXiv:2604.05007cs.SDcs.AI2026-04中稿 · publication by the…

通过双耳差异注意力与动作过渡预测提升音频视觉导航泛化能力

Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction

  • 引入双耳差异注意力模块,增强空间定位能力
  • 通过动作过渡预测任务提升在未见环境中的表现,最高提升21.6个百分点
  • 适用于多种导航架构,特别适合需要强泛化的音频视觉任务

在音频-视觉导航(AVN)中,智能体需在未见过的3D环境中利用视觉和听觉线索定位声源。现有方法常因过度依赖语义声音特征和特定训练环境而泛化能力不足。为此,我们提出双耳差异注意力与动作过渡预测(BDATP)框架,联合优化感知与策略。其中,双耳差异注意力(BDA)模块显式建模双耳间差异,提升空间方向感知,降低对语义类别依赖;动作过渡预测(ATP)作为辅助目标引入正则化项,缓解环境特异性过拟合。在Replica和Matterport3D数据集上的大量实验表明,BDATP可无缝集成至多种主流基线,实现一致且显著的性能提升。尤其在未听过的声音上,Replica数据集上成功率最高提升21.6个百分点。结果验证了该框架卓越的泛化能力及对不同导航架构的鲁棒性。

原文摘要 · Abstract (English)

In Audio-Visual Navigation (AVN), agents must locate sound sources in unseen 3D environments using visual and auditory cues. However, existing methods often struggle with generalization in unseen scenarios, as they tend to overfit to semantic sound features and specific training environments. To address these challenges, we propose the \textbf{Binaural Difference Attention with Action Transition Prediction (BDATP)} framework, which jointly optimizes perception and policy. Specifically, the \textbf{Binaural Difference Attention (BDA)} module explicitly models interaural differences to enhance spatial orientation, reducing reliance on semantic categories. Simultaneously, the \textbf{Action Transition Prediction (ATP)} task introduces an auxiliary action prediction objective as a regularization term, mitigating environment-specific overfitting. Extensive experiments on the Replica and Matterport3D datasets demonstrate that BDATP can be seamlessly integrated into various mainstream baselines, yielding consistent and significant performance gains. Notably, our framework achieves state-of-the-art Success Rates across most settings, with a remarkable absolute improvement of up to 21.6 percentage points in Replica dataset for unheard sounds. These results underscore BDATP's superior generalization capability and its robustness across diverse navigation architectures.

音频视觉导航空间定位泛化能力双耳注意

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。