让智能体在连续3D空间中靠声视觉导航,应对目标间歇性沉默挑战。
Semantic Audio-Visual Navigation in Continuous Environments
- 用多模态变换器融合空间与语义目标信息,结合历史上下文和自身运动线索。
- 在短时声响和远距离导航场景下,成功率提升12.1%。
- 适合研究真实环境下的声视觉导航与记忆增强决策的学者。
声视觉导航使具身智能体通过听觉与视觉线索向发声目标移动。然而,现有方法多依赖预计算的房间冲激响应(RIR)进行双耳音频渲染,导致智能体仅能停留在离散网格位置,观察结果存在空间不连续性。为建立更真实的场景,我们提出连续环境中的语义声视觉导航(SAVN-CE),允许智能体在3D空间自由移动,并感知时空连贯的声视觉流。在此设置中,目标可能间歇性静音或完全停止发声,导致智能体丢失目标信息。为此,我们提出MAGNet,一种基于多模态变换器的模型,联合编码空间与语义目标表征,融合历史上下文与自运动线索,实现增强记忆的目标推理。全面实验表明,MAGNet显著优于当前最优方法,在成功率上最高提升12.1%。结果还验证了其对短时声音及长距离导航场景的鲁棒性。代码已开源:https://github.com/yichenzeng24/SAVN-CE。
原文摘要 · Abstract (English)
Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio rendering, restricting agents to discrete grid positions and leading to spatially discontinuous observations. To establish a more realistic setting, we introduce Semantic Audio-Visual Navigation in Continuous Environments (SAVN-CE), where agents can move freely in 3D spaces and perceive temporally and spatially coherent audio-visual streams. In this setting, targets may intermittently become silent or stop emitting sound entirely, causing agents to lose goal information. To tackle this challenge, we propose MAGNet, a multimodal transformer-based model that jointly encodes spatial and semantic goal representations and integrates historical context with self-motion cues to enable memory-augmented goal reasoning. Comprehensive experiments demonstrate that MAGNet significantly outperforms state-of-the-art methods, achieving up to a 12.1\% absolute improvement in success rate. These results also highlight its robustness to short-duration sounds and long-distance navigation scenarios. The code is available at https://github.com/yichenzeng24/SAVN-CE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。