用视觉语言模型和世界模型结合,让机器人在新场景中自主规划动作。
World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models

- 利用视觉语言模型生成动作计划,并通过世界模型模拟优化
- 在组合任务和零样本泛化上超越现有端到端模型
- 适合需要跨场景决策的机器人系统研究者
构建适用于多样化应用的通用智能体仍是核心挑战。尽管基于模仿学习的策略在特定训练环境中表现良好,但往往难以推广到新场景与新任务。本文提出 World Action Planner,一种融合视觉语言模型推理能力与多任务位姿-图像条件世界模型物理基础的机器人规划系统。该系统使智能体能够生成初始动作计划,并通过优化与搜索迭代修正,基于想象的世界模型回溯进行推理。实验表明,该方法在组合任务、新布局及零样本泛化场景中均表现出色,显著优于当前最先进的端到端策略模型(如 VLAs 与 WAMs)。项目网站:worldactionplanner.github.io
原文摘要 · Abstract (English)
Building generalizable agents for diverse applications remains a fundamental challenge. While imitation learning-based policies succeed in specific training environments, they often fail to generalize to novel scenes and tasks. In this work, we propose World Action Planner, a robot planning system that leverages the reasoning capabilities of Vision-Language Models (VLMs) and the physical grounding of a multi-task pose-image conditioned world model. Our system enables an agent to propose initial action plans and iteratively refine them via optimization and search, reasoning over imagined world model rollouts. We demonstrate that our approach achieves superior performance across compositional tasks, new layouts, and zero-shot generalization scenarios, significantly outperforming state-of-the-art end-to-end policy models such as VLAs and WAMs. Project website at worldactionplanner.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。