arXiv:2504.03245cs.AIcs.RO2025-04

用视觉语言模型估计不确定性,让机器人在信息不足时也能智能规划。

Seeing is Believing: Belief-Space Planning with Foundation Models as Uncertainty Estimators

  • 用VLM构建符号化信念表示,动态评估环境不确定性
  • 在部分可观测场景中规划并执行主动获取信息的动作
  • 适合需要应对信息不全的复杂现实任务的机器人系统

开放世界中的通用机器人移动操作面临长时程、复杂目标和部分可观测性的挑战。现有方法常依赖参数化技能库与任务规划器组合实现目标,通过逻辑表达式等结构化语言定义目标。尽管视觉语言模型(VLMs)可用于语义接地,但其通常假设完全可观测,导致信息不足时行为次优。本文提出新框架,将VLM作为感知模块,用于估计不确定性并支持符号化接地。该方法构建符号化信念表示,并采用信念空间规划器生成包含主动信息获取策略的计划,使智能体能有效处理部分可观测性与属性不确定性。我们在一系列需在部分可观测环境中推理的真实世界任务上验证系统性能。仿真结果表明,相比端到端的VLM规划或基于VLM的状态估计基线,本方法通过规划并执行战略信息收集显著提升表现。该工作展示了VLM构建信念空间符号化场景表征的潜力,为不确定性感知规划等下游任务提供支持。

原文摘要 · Abstract (English)

Generalizable robotic mobile manipulation in open-world environments poses significant challenges due to long horizons, complex goals, and partial observability. A promising approach to address these challenges involves planning with a library of parameterized skills, where a task planner sequences these skills to achieve goals specified in structured languages, such as logical expressions over symbolic facts. While vision-language models (VLMs) can be used to ground these expressions, they often assume full observability, leading to suboptimal behavior when the agent lacks sufficient information to evaluate facts with certainty. This paper introduces a novel framework that leverages VLMs as a perception module to estimate uncertainty and facilitate symbolic grounding. Our approach constructs a symbolic belief representation and uses a belief-space planner to generate uncertainty-aware plans that incorporate strategic information gathering. This enables the agent to effectively reason about partial observability and property uncertainty. We demonstrate our system on a range of challenging real-world tasks that require reasoning in partially observable environments. Simulated evaluations show that our approach outperforms both vanilla VLM-based end-to-end planning or VLM-based state estimation baselines by planning for and executing strategic information gathering. This work highlights the potential of VLMs to construct belief-space symbolic scene representations, enabling downstream tasks such as uncertainty-aware planning.

机器人不确定性视觉语言模型规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。