arXiv:2605.23165cs.ROcs.AI2026-05

用视觉语言模型指导机器人探索未知环境,提升地图覆盖效率。

Autonomous Frontier-Based Exploration with VLM Guidance

论文配图:Autonomous Frontier-Based Exploration with VLM Guidance
图 1 · 摘自论文原文
  • 用视觉语言模型分析多模态提示,智能选择最佳探索路径。
  • 在六种室内环境中,地图覆盖率最高提升24%。
  • 无需训练,适配普通机器人,适合快速部署的探索任务。

自主机器人在未知且危险环境中的探索是一个长期挑战,可通过利用视觉语言模型(VLM)的先进推理能力显著改进。我们提出一种新型探索流程,其中VLM负责高层次战略决策,指导传统的低层机器人控制栈。在决策点,机器人结合当前地图和潜在路径的视觉图像生成多模态提示,VLM分析该提示后选择最具前景的前沿区域,以情境化空间推理替代简单的几何启发式方法。该方法在六种室内环境的仿真中验证,相较于现有方法,地图覆盖率最高提升24%。本流程轻量、免训练,可轻松迁移至配备标准传感器和互联网连接的任意机器人。

原文摘要 · Abstract (English)

Autonomous robotic exploration of unknown and hazardous environments, a long-standing challenge, can be significantly improved by leveraging the advanced reasoning of Vision-Language Models (VLMs). We introduce a novel exploration pipeline where a VLM performs high-level strategic decision-making, guiding a conventional low-level robotics control stack. At decision points, the robot generates a multimodal prompt with its current map and visual imagery of potential paths, or frontiers. The VLM analyzes this prompt to select the most promising frontier, replacing simple geometric heuristics with contextual spatial reasoning. This approach, validated in simulation across six indoor environments, improves map coverage by up to 24\% over existing methods. Our pipeline is lightweight, training-free, and easily transferable to any robot with standard sensors and an internet connection.

自主探索视觉语言模型机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。