arXiv:2505.07815cs.ROcs.CV2025-05被引 2

用视觉语言模型实现有想象、可验证的机器人探索,提升环境学习效率。

Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models

  • 基于视觉语言模型构建语义场景图,生成并验证想象中的新场景。
  • 探索状态熵提升4.1到7.8倍,显著增加探索多样性。
  • 适合需要自主探索与少样本学习的机器人系统研究者。

探索对通用机器人学习至关重要,尤其在奖励稀疏、目标不明确的开放环境。视觉语言模型(VLM)具备对物体、空间关系及潜在结果的语义推理能力,可作为生成高层探索行为的基础。然而其输出常缺乏物理依据,难以判断想象是否可行。为此,我们提出IVE(Imagine, Verify, Execute)框架,模仿人类好奇心驱动的探索模式。该框架将RGB-D观测抽象为语义场景图,想象新场景,预测其物理合理性,并通过动作工具生成可执行技能序列。在仿真与真实桌面环境中评估显示,相比强化学习基线,IVE使访问状态的熵提高4.1至7.8倍,探索更丰富且有意义。所收集经验有效支持下游任务学习,训练出的策略性能接近或超过人类示范数据训练的结果。

原文摘要 · Abstract (English)

Exploration is essential for general-purpose robotic learning, especially in open-ended environments where dense rewards, explicit goals, or task-specific supervision are scarce. Vision-language models (VLMs), with their semantic reasoning over objects, spatial relations, and potential outcomes, present a compelling foundation for generating high-level exploratory behaviors. However, their outputs are often ungrounded, making it difficult to determine whether imagined transitions are physically feasible or informative. To bridge the gap between imagination and execution, we present IVE (Imagine, Verify, Execute), an agentic exploration framework inspired by human curiosity. Human exploration is often driven by the desire to discover novel scene configurations and to deepen understanding of the environment. Similarly, IVE leverages VLMs to abstract RGB-D observations into semantic scene graphs, imagine novel scenes, predict their physical plausibility, and generate executable skill sequences through action tools. We evaluate IVE in both simulated and real-world tabletop environments. The results show that IVE enables more diverse and meaningful exploration than RL baselines, as evidenced by a 4.1 to 7.8x increase in the entropy of visited states. Moreover, the collected experience supports downstream learning, producing policies that closely match or exceed the performance of those trained on human-collected demonstrations.

机器人探索视觉语言模型自适应学习智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。