arXiv:2505.13180cs.AI2025-05被引 8

首个视觉规划基准,对比视觉语言模型作规划器与地标的优劣

ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models

  • 构建双场景基准:积木世界与家庭机器人模拟环境
  • 地标的模型在积木任务中解决46%任务,规划器在家庭任务中解决34%
  • 揭示视觉推理局限,适合研究多模态规划与评估的学者

将大语言模型与符号规划器结合是实现可验证、具身化规划的有前景方向,近期工作已将其拓展至使用视觉语言模型(VLMs)的视觉领域。然而,由于缺乏支持符号规划的开源视觉基准,现有方法难以在一致条件下比较。本文提出ViPlan,首个开放源代码的基准,用于对比VLM作为地标的规划方法(VLM-as-grounder)与直接使用VLM进行规划的方法(VLM-as-planner)。ViPlan包含两个视觉领域的递进挑战任务:经典积木世界问题的视觉变体和模拟家庭机器人环境。在积木任务中,平均而言,地标的模型解决46%的任务,而直接规划方法仅解决9%,因图像定位至关重要且准确。但在家庭机器人任务中,依赖语言知识的规划器方法显著优于地标方法(34%对5%),后者受部分可观测性限制。因此,ViPlan揭示了两类方法的根本缺陷,通过定性失败分析进一步诊断。此外,所有方法均未表现出链式思维提示的一致优势,表明当前VLM在视觉推理方面仍存在持久局限。

原文摘要 · Abstract (English)

Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent works extending this idea to visual domains using Vision-Language Models (VLMs). However, an open-source benchmark for comparing these approaches under matched conditions is missing, due to a lack of visual benchmarks that support symbolic planning. We present ViPlan, the first open-source benchmark for comparing VLM-grounded symbolic approaches (VLM-as-grounder) with direct VLM planning methods (VLM-as-planner). ViPlan introduces a series of increasingly challenging tasks in two visual domains: a visual variant of the classic Blocksworld planning problem and a simulated household robotics environment. Averaged across methods, we find VLM-as-grounders to outperform direct VLM planning in Blocksworld (solving 46% of the tasks against 9%), where image grounding is both crucial and accurate. However, in the household robotics tasks, where linguistic knowledge helps, VLM-as-planner methods are greatly superior to VLM-as-grounder approaches (solving 34% of the tasks against 5%), which are hindered by partial observability. Thus, ViPlan domains capture fundamental shortcomings of both planning approaches, which we further diagnose with a qualitative failure analysis. Finally, across methods, we observe no consistent benefit from Chain-of-Thought prompting, suggesting persistent limitations in current VLMs' visual reasoning abilities.

视觉规划多模态基准测试语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。