提出新型端到端机制,提升音视频导航精度与稳定性。
Iterative Residual Cross-Attention Mechanism: An Integrated Approach for Audio-Visual Navigation Tasks
- 用迭代残差交叉注意力融合视听信息,替代传统分步模块。
- 在Audio-Visual Navigation任务中实现更优导航性能。
- 适合研究多模态感知与智能体决策的学者参考。
音视频导航是智能体利用第一人称视觉与听觉感知定位声源的重要研究方向。传统方法通常采用分阶段模块化设计:先进行特征融合,再通过门控循环单元(GRU)建模序列,最后通过强化学习做出决策。尽管有效,但该方式易导致信息冗余及模块间传递不一致。本文提出IRCAM-AVN(用于音视频导航的迭代残差交叉注意力机制),一种端到端框架,将多模态特征融合与序列建模统一整合至一个IRCAM模块中,取代传统的独立融合与GRU组件。该机制采用多层次残差设计,将原始多模态序列与处理后的信息序列拼接,逐步优化特征提取过程,降低模型偏差,增强模型稳定性和泛化能力。实验表明,采用该机制的智能体在音视频导航任务中表现更优。
原文摘要 · Abstract (English)
Audio-visual navigation represents a significant area of research in which intelligent agents utilize egocentric visual and auditory perceptions to identify audio targets. Conventional navigation methodologies typically adopt a staged modular design, which involves first executing feature fusion, then utilizing Gated Recurrent Unit (GRU) modules for sequence modeling, and finally making decisions through reinforcement learning. While this modular approach has demonstrated effectiveness, it may also lead to redundant information processing and inconsistencies in information transmission between the various modules during the feature fusion and GRU sequence modeling phases. This paper presents IRCAM-AVN (Iterative Residual Cross-Attention Mechanism for Audiovisual Navigation), an end-to-end framework that integrates multimodal information fusion and sequence modeling within a unified IRCAM module, thereby replacing the traditional separate components for fusion and GRU. This innovative mechanism employs a multi-level residual design that concatenates initial multimodal sequences with processed information sequences. This methodological shift progressively optimizes the feature extraction process while reducing model bias and enhancing the model's stability and generalization capabilities. Empirical results indicate that intelligent agents employing the iterative residual cross-attention mechanism exhibit superior navigation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。