arXiv:2411.05831cs.AIcs.CV2024-11中稿 · WACV 2025

让视觉语言导航智能体学会识别指令模糊时刻,主动提问提升效率。

To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation

  • 用注意力机制建模指令与路径对齐,判断何时信息不足。
  • 精度与召回平衡提升52%,比传统方法更准地识别模糊点。
  • 适合需要自主决策的智能体导航场景,如机器人导览。

当前视觉语言导航(VLN)研究忽视了智能体的提问能力,即在指令不完整时主动询问澄清信息。本文聚焦于如何让智能体识别‘何时’缺乏足够信息,而不关心‘缺什么’,尤其针对含糊指令的任务。赋予智能体这种能力可减少无效探索,及时求助,提升效率。识别不确定性的难点在于平衡过度谨慎(高召回)与过度自信(高精度)。我们提出一种基于注意力的指令模糊度估计模块,学习指令与智能体轨迹之间的关联。通过训练中利用指令-路径对齐信息,该模块的模糊度估计性能在精度-召回平衡上提升了约52%。消融实验表明,将此指令-路径注意力网络与导航器中的跨模态注意力网络结合使用效果更佳。结果发现,该注意力网络的得分能更优地指示指令模糊程度。

原文摘要 · Abstract (English)

Recent research in Vision Language Navigation (VLN) has overlooked the development of agents' inquisitive abilities, which allow them to ask clarifying questions when instructions are incomplete. This paper addresses how agents can recognize "when" they lack sufficient information, without focusing on "what" is missing, particularly in VLN tasks with vague instructions. Equipping agents with this ability enhances efficiency by reducing potential digressions and seeking timely assistance. The challenge in identifying such uncertain points is balancing between being overly cautious (high recall) and overly confident (high precision). We propose an attention-based instruction-vagueness estimation module that learns associations between instructions and the agent's trajectory. By leveraging instruction-to-path alignment information during training, the module's vagueness estimation performance improves by around 52% in terms of precision-recall balance. In our ablative experiments, we also demonstrate the effectiveness of incorporating this additional instruction-to-path attention network alongside the cross-modal attention networks within the navigator module. Our results show that the attention scores from the instruction-to-path attention network serve as better indicators for estimating vagueness.

视觉语言导航智能体决策注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。