arXiv:2509.20499cs.ROcs.AI2025-09被引 4

用抽象障碍图与拓扑图提示,提升零样本视觉语言导航的探索能力

Boosting Zero-Shot VLN via Abstract Obstacle Map-Based Waypoint Prediction with TopoGraph-and-VisitInfo-Aware Prompting

  • 基于抽象障碍图生成可到达路径点,融合拓扑结构与访问记录
  • 在R2R-CE和RxR-CE上实现41%和36%的零样本成功率
  • 适合需要低资源部署的机器人导航任务

随着基础模型与机器人技术的快速发展,视觉语言导航(VLN)已成为具身智能体的关键任务,具有广泛的应用前景。本文聚焦连续环境下的零样本视觉语言导航,该场景要求智能体联合理解自然语言指令、感知周围环境并规划底层动作。我们提出一种零样本框架,结合简化而高效的路径点预测器与多模态大语言模型(MLLM)。预测器基于抽象障碍图生成线性可达的路径点,并将其融入动态更新的拓扑图,显式记录访问历史。拓扑图与访问信息被编码进提示中,使模型能推理空间结构与探索历史,促进探索行为,并赋予MLLM局部路径规划能力以纠正错误。在R2R-CE和RxR-CE上的大量实验表明,本方法达到当前最优的零样本性能,成功率达41%和36%,优于已有最先进方法。

原文摘要 · Abstract (English)

With the rapid progress of foundation models and robotics, vision-language navigation (VLN) has emerged as a key task for embodied agents with broad practical applications. We address VLN in continuous environments, a particularly challenging setting where an agent must jointly interpret natural language instructions, perceive its surroundings, and plan low-level actions. We propose a zero-shot framework that integrates a simplified yet effective waypoint predictor with a multimodal large language model (MLLM). The predictor operates on an abstract obstacle map, producing linearly reachable waypoints, which are incorporated into a dynamically updated topological graph with explicit visitation records. The graph and visitation information are encoded into the prompt, enabling reasoning over both spatial structure and exploration history to encourage exploration and equip MLLM with local path planning for error correction. Extensive experiments on R2R-CE and RxR-CE show that our method achieves state-of-the-art zero-shot performance, with success rates of 41% and 36%, respectively, outperforming prior state-of-the-art methods.

视觉语言导航零样本路径规划大语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。