用视觉语言模型实现机器人闭环探索与任务规划
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
- 基于双阶段自省式规划与空间关系图,实现动态环境中的感知与决策
- 实测表明在探索类任务中显著优于现有方法,成功率更高
- 适合需要交互探索与实时调整的智能机器人应用
具身智能的发展正推动机器人作为人类助手融入日常生活。这要求机器人不仅能理解高层指令并规划任务,还需在动态环境中感知与适应。视觉语言模型(VLMs)通过结合视觉理解与语言推理,展现出巨大潜力。然而,现有基于VLM的方法在交互式探索、精确感知和实时计划调整方面仍存在不足。为此,我们提出ExploreVLM,一种由视觉语言模型驱动的新型闭环任务规划框架。该框架采用分步反馈机制,支持实时计划调整与交互探索。核心是一个具有自省能力的双阶段任务规划器,结合以对象为中心的空间关系图,提供结构化、语言对齐的场景表示,指导感知与规划。执行验证模块通过验证每一步动作并触发重规划,构建闭环。大量真实世界实验表明,ExploreVLM在探索导向任务中显著优于现有先进基准。消融实验进一步验证了自省式规划器与结构化感知在实现鲁棒高效任务执行中的关键作用。
原文摘要 · Abstract (English)
The advancement of embodied intelligence is accelerating the integration of robots into daily life as human assistants. This evolution requires robots to not only interpret high-level instructions and plan tasks but also perceive and adapt within dynamic environments. Vision-Language Models (VLMs) present a promising solution by combining visual understanding and language reasoning. However, existing VLM-based methods struggle with interactive exploration, accurate perception, and real-time plan adaptation. To address these challenges, we propose ExploreVLM, a novel closed-loop task planning framework powered by Vision-Language Models (VLMs). The framework is built around a step-wise feedback mechanism that enables real-time plan adjustment and supports interactive exploration. At its core is a dual-stage task planner with self-reflection, enhanced by an object-centric spatial relation graph that provides structured, language-grounded scene representations to guide perception and planning. An execution validator supports the closed loop by verifying each action and triggering re-planning. Extensive real-world experiments demonstrate that ExploreVLM significantly outperforms state-of-the-art baselines, particularly in exploration-centric tasks. Ablation studies further validate the critical role of the reflective planner and structured perception in achieving robust and efficient task execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。