arXiv:2512.00027cs.ROcs.AI2025-12综述被引 1

综述视觉语言导航在人机协作中的进展与挑战

A Survey on Improving Human Robot Collaboration through Vision-and-Language Navigation

论文配图:A Survey on Improving Human Robot Collaboration through Vision-and-Language Navigation
图 1 · 摘自论文原文
  • 梳理200篇文献,分析视觉语言导航在机器人中的应用
  • 指出当前模型在双向沟通和协同决策上存在不足
  • 适合关注人机交互与多机器人协作的研究者

视觉-语言导航(VLN)是一项多模态协作任务,要求智能体理解人类指令、在三维环境中导航并有效沟通以应对歧义。本文系统回顾了近年来机器人领域中VLN的进展,并展望了提升多机器人协调的潜在方向。尽管已有一定进展,当前模型仍难以实现双向通信、歧义解析以及多智能体系统中的协同决策。我们分析了约200篇相关论文,旨在深入理解该领域的现状。本综述强调,未来VLN系统应通过先进的自然语言理解(NLU)技术实现主动澄清、实时反馈和上下文推理。此外,具备动态角色分配的去中心化决策框架对实现可扩展、高效的多机器人协作至关重要。这些创新有望显著提升人机交互(HRI)能力,推动其在医疗、物流及灾难救援等现实场景中的应用。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) is a multi-modal, cooperative task requiring agents to interpret human instructions, navigate 3D environments, and communicate effectively under ambiguity. This paper presents a comprehensive review of recent VLN advancements in robotics and outlines promising directions to improve multi-robot coordination. Despite progress, current models struggle with bidirectional communication, ambiguity resolution, and collaborative decision-making in the multi-agent systems. We review approximately 200 relevant articles to provide an in-depth understanding of the current landscape. Through this survey, we aim to provide a thorough resource that inspires further research at the intersection of VLN and robotics. We advocate that the future VLN systems should support proactive clarification, real-time feedback, and contextual reasoning through advanced natural language understanding (NLU) techniques. Additionally, decentralized decision-making frameworks with dynamic role assignment are essential for scalable, efficient multi-robot collaboration. These innovations can significantly enhance human-robot interaction (HRI) and enable real-world deployment in domains such as healthcare, logistics, and disaster response.

视觉语言导航人机协作多机器人自然语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。