arXiv:2607.03920cs.RO2026-07

构建首个支持多目标、跨模态指令的长时音频视觉导航基准

LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation

论文配图:LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation
图 1 · 摘自论文原文
  • 设计多目标、混合指令的长时导航任务,融合语音与视觉线索
  • 现有模型在复杂任务中完成率不足30%,暴露出泛化缺陷
  • 适合研究多模态融合、长期规划与听觉导航的学者使用

具身导航正向长时程任务演进,但现有基准多为无声环境,且通常仅关注单一目标。本文提出LH-AVLN,一个支持多目标、异构目标描述和持续空间化声学提示的长时程音频-视觉-语言导航基准。在该基准中,智能体接收包含2至4个目标的全局任务,目标可由类别、语言描述或参考图像指定,通过RGB-D观测、位姿信息和双耳音频在室内3D环境中导航。任务支持有序与无序执行模式,交替出现的目标相关声音可引导非视距搜索,但也可能随任务进展成为干扰项。我们进一步提出PAG-Nav——一种无需训练的参考智能体,通过维护时间一致的语义地图,实现渐进式目标状态规划,利用声音进行搜索,同时保留视觉语义验证以完成任务。实验表明,现有视觉-语言、记忆型及音视频模型难以完整完成LH-AVLN任务,而PAG-Nav提供了更强的诊断基线,仍留有巨大提升空间。

原文摘要 · Abstract (English)

Embodied navigation is moving toward long-horizon missions, yet existing long-horizon benchmarks are largely acoustically silent, and audio-visual navigation tasks typically focus on a single goal. We introduce LH-AVLN, a benchmark for Long-Horizon Audio-Visual-Language Navigation that combines multi-goal mission execution, heterogeneous goal specifications, and persistent spatialized acoustic cues. In LH-AVLN, an agent receives a global mission of two to four goals specified by category, language description, or reference image, and navigates with RGB-D observations, pose, and binaural audio in indoor 3D environments. The benchmark supports both ordered and unordered missions, where alternating goal-associated sounds can guide non-line-of-sight search but may also become distractors as mission progress changes. We further develop PAG-Nav, a training-free reference agent that maintains a temporal uniform semantic map and performs progressive goal-state planning, using sound for search while reserving completion for visual-semantic verification. Experiments show that existing vision-language, memory-based, and audio-visual agents struggle to complete full LH-AVLN missions, and that PAG-Nav provides a stronger diagnostic baseline while leaving substantial room for future progress.

导航基准多模态长时程音频感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。