多智能体协作让声音导航更快更准,适合应急搜救等紧急场景。
Advancing Audio-Visual Navigation Through Multi-Agent Collaboration in 3D Environments
- 两智能体通过通信与视听融合协同定位声音目标
- 任务完成时间减少,导航成功率显著高于单智能体方案
- 适用于需要快速响应的3D复杂环境,如灾后搜救
智能体在现实世界中常需协作才能完成个体无法胜任的复杂任务。现有音频-视觉导航(AVN)研究主要聚焦单智能体系统,但在动态3D环境中,其局限性凸显,尤其在应急响应等对时间敏感的应用中。本文提出MASTAVN(多智能体可扩展变压器音频-视觉导航),一种支持两个智能体在共享3D环境中协作定位并导航至声源的可扩展框架。通过引入跨智能体通信协议与联合视听融合机制,MASTAVN增强了空间推理与时间同步能力。在逼真的3D模拟器Replica和Matterport3D中的严格评估表明,相比单智能体及非协作基线,MASTAVN在任务完成时间上显著降低,导航成功率明显提升。这验证了时空协调在多智能体系统中的关键作用。研究结果证明MASTAVN在时间敏感的应急场景中有效,并为复杂3D环境中可扩展多智能体具身智能的发展提供了新范式。
原文摘要 · Abstract (English)
Intelligent agents often require collaborative strategies to achieve complex tasks beyond individual capabilities in real-world scenarios. While existing audio-visual navigation (AVN) research mainly focuses on single-agent systems, their limitations emerge in dynamic 3D environments where rapid multi-agent coordination is critical, especially for time-sensitive applications like emergency response. This paper introduces MASTAVN (Multi-Agent Scalable Transformer Audio-Visual Navigation), a scalable framework enabling two agents to collaboratively localize and navigate toward an audio target in shared 3D environments. By integrating cross-agent communication protocols and joint audio-visual fusion mechanisms, MASTAVN enhances spatial reasoning and temporal synchronization. Through rigorous evaluation in photorealistic 3D simulators (Replica and Matterport3D), MASTAVN achieves significant reductions in task completion time and notable improvements in navigation success rates compared to single-agent and non-collaborative baselines. This highlights the essential role of spatiotemporal coordination in multi-agent systems. Our findings validate MASTAVN's effectiveness in time-sensitive emergency scenarios and establish a paradigm for advancing scalable multi-agent embodied intelligence in complex 3D environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。