用预演视觉引导语言动作模型,提升复杂任务成功率。
ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
- 通过生成未来视觉画面,让模型专注执行而非思考语义。
- 0.33秒内生成640×480高质未来图像,平均成功率达87.4%。
- 无需修改模型结构,适配现有视觉语言动作系统。
视觉-语言-动作(VLA)模型将高层语言指令转化为可执行动作,尤其在开放世界中极具挑战。本文提出视觉预演规划(ForeAct),一种通用且高效的规划器,通过想象的未来观测和子任务描述逐步引导VLA。借助想象的未来视觉输入,VLA可专注于视觉运动推理,而非高层语义理解,从而提升准确率与泛化能力。该规划器包含一个高效预演图像生成模块,在H100 GPU上仅需0.33秒即可从当前视觉输入与语言指令生成640×480高质量未来图像;同时配备视觉语言模型,用于任务推理并生成子任务描述以指导生成器与VLA。重要的是,主流VLA可通过简单扩展视觉输入无缝集成本框架,无需架构修改。预演生成器在超过100万条跨任务、跨体感的训练数据上预训练,具备鲁棒的具身动力学建模能力。我们在包含11个多样化多步真实任务的基准上评估,平均成功率达87.4%,相较π₀基线(46.5%)提升40.9个百分点,较添加文本子任务引导的π₀(57.1%)提升30.3个百分点。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models convert high-level language instructions into concrete, executable actions, a task that is especially challenging in open-world environments. We present Visual Foresight Planning (ForeAct), a general and efficient planner that guides a VLA step-by-step using imagined future observations and subtask descriptions. With an imagined future observation, the VLA can focus on visuo-motor inference rather than high-level semantic reasoning, leading to improved accuracy and generalization. Our planner comprises a highly efficient foresight image generation module that predicts a high-quality 640$\times$480 future observation from the current visual input and language instruction within only 0.33s on an H100 GPU, together with a vision-language model that reasons over the task and produces subtask descriptions for both the generator and the VLA. Importantly, state-of-the-art VLAs can integrate our planner seamlessly by simply augmenting their visual inputs, without any architectural modification. The foresight generator is pretrained on over 1 million multi-task, cross-embodiment episodes, enabling it to learn robust embodied dynamics. We evaluate our framework on a benchmark that consists of 11 diverse, multi-step real-world tasks. It achieves an average success rate of 87.4%, demonstrating a +40.9% absolute improvement over the $π_0$ baseline (46.5%) and a +30.3% absolute improvement over $π_0$ augmented with textual subtask guidance (57.1%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。