提出残差交叉模态融合网络,提升音视频导航的跨域泛化能力。
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
- 通过双向残差连接实现音视频模态间的互补建模与精细对齐
- 在Replica和Matterport3D上显著超越现有基线,跨域泛化更强
- 揭示不同数据集下代理对模态依赖性的差异,为多模态协作提供新视角
音视频具身导航旨在使智能体利用听觉线索,在未见过的3D环境中自主定位并抵达声源。该任务的关键挑战在于如何有效建模异构特征在多模态融合过程中的交互,避免单一模态主导或信息退化,尤其是在跨领域场景中。为此,本文提出交叉模态残差融合网络(CRFN),通过音视频流之间的双向残差交互,实现互补建模与细粒度对齐,同时保持各自表示的独立性。与依赖简单拼接或注意力门控的传统方法不同,CRFN通过残差连接显式建模跨模态交互,并引入稳定化技术以提升收敛性和鲁棒性。在Replica和Matterport3D数据集上的实验表明,CRFN显著优于现有先进融合基线,且具备更强的跨域泛化能力。值得注意的是,实验还发现智能体在不同数据集上表现出不同的模态依赖性。这一现象为理解具身智能体的跨模态协作机制提供了新视角。
原文摘要 · Abstract (English)
Audio-visual embodied navigation aims to enable an agent to autonomously localize and reach a sound source in unseen 3D environments by leveraging auditory cues. The key challenge of this task lies in effectively modeling the interaction between heterogeneous features during multimodal fusion, so as to avoid single-modality dominance or information degradation, particularly in cross-domain scenarios. To address this, we propose a Cross-Modal Residual Fusion Network, which introduces bidirectional residual interactions between audio and visual streams to achieve complementary modeling and fine-grained alignment, while maintaining the independence of their representations. Unlike conventional methods that rely on simple concatenation or attention gating, CRFN explicitly models cross-modal interactions via residual connections and incorporates stabilization techniques to improve convergence and robustness. Experiments on the Replica and Matterport3D datasets demonstrate that CRFN significantly outperforms state-of-the-art fusion baselines and achieves stronger cross-domain generalization. Notably, our experiments also reveal that agents exhibit differentiated modality dependence across different datasets. The discovery of this phenomenon provides a new perspective for understanding the cross-modal collaboration mechanism of embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。