arXiv:2506.21230cs.AIcs.RO2025-06NeurIPS被引 11

让视觉语言模型理解环境,提升复杂任务规划能力。

World-aware Planning Narratives Enhance Large Vision-Language Model Planner

  • 通过视觉、空间、功能和语法四方面增强模型对环境的认知。
  • 在EB-ALFRED上任务成功率提升60.7,长程规划提升70.0。
  • 开源模型性能超越GPT-4o和Claude-3.5-Sonnet,适合实际部署。

大型视觉语言模型(LVLMs)在具身规划任务中展现潜力,但在涉及陌生环境和多步骤目标的复杂场景中表现不佳。现有方法依赖与环境无关的模仿学习,使指令脱离上下文,导致模型难以处理情境相关指令,需依赖额外提示而非视觉推理。本文提出世界感知规划叙事增强(WAP)框架,通过视觉外观建模、空间推理、功能抽象和句法定位四类认知能力,赋予LVLM全面的环境理解力,并仅使用原始视觉观察进行课程学习训练。在EB-ALFRED基准上的评估显示显著提升:Qwen2.5-VL任务成功率提高60.7,常识推理提升60.0,长程规划提升70.0。值得注意的是,我们推出的开源模型大幅超越专有系统如GPT-4o和Claude-3.5-Sonnet。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) show promise for embodied planning tasks but struggle with complex scenarios involving unfamiliar environments and multi-step goals. Current approaches rely on environment-agnostic imitation learning that disconnects instructions from environmental contexts, causing models to struggle with context-sensitive instructions and rely on supplementary cues rather than visual reasoning during long-horizon interactions. In this work, we propose World-Aware Planning Narrative Enhancement (WAP), a framework that infuses LVLMs with comprehensive environmental understanding through four cognitive capabilities (visual appearance modeling, spatial reasoning, functional abstraction, and syntactic grounding) while developing and evaluating models using only raw visual observations through curriculum learning. Evaluations on the EB-ALFRED benchmark demonstrate substantial improvements, with Qwen2.5-VL achieving a 60.7 absolute improvement in task success rates, particularly in commonsense reasoning (+60.0) and long-horizon planning (+70.0). Notably, our enhanced open-source models outperform proprietary systems like GPT-4o and Claude-3.5-Sonnet by a large margin.

视觉语言模型具身智能规划生成环境理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。