让机器只用图像思考,提升视觉任务的规划能力
Visual Planning: Let's Think Only with Images
- 用图像序列代替文字进行视觉推理规划
- 在多个导航任务中性能超越纯文本规划方法
- 适合需要直观空间推理的视觉任务研究者
近年来,大语言模型及其多模态扩展显著提升了机器在多样化任务中的推理能力。然而,这些模型在处理包含视觉信息的任务时,仍主要依赖纯文本进行表达和结构化推理。本文提出一种新范式——视觉规划(Visual Planning),主张在涉及空间与几何信息的任务中,图像可能比语言更自然高效。该范式通过一系列图像序列实现视觉领域的逐步推理,类似人类画草图预演动作。我们引入基于GRPO的强化学习框架VPRL,对大型视觉模型进行后训练,在FrozenLake、Maze和MiniBehavior等典型视觉导航任务上取得显著提升。结果表明,视觉规划在性能上优于所有仅使用文本推理的规划方法,验证了其作为语言推理补充的可行性与前景,为依赖直觉图像推理的任务开辟新路径。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across diverse tasks. However, these models predominantly rely on pure text as the medium for both expressing and structuring reasoning, even when visual information is present. In this work, we argue that language may not always be the most natural or effective modality for reasoning, particularly in tasks involving spatial and geometrical information. Motivated by this, we propose a new paradigm, Visual Planning, which enables planning through purely visual representations for these "vision-first" tasks, as a supplementary channel to language-based reasoning. In this paradigm, planning is executed via sequences of images that encode step-by-step inference in the visual domain, akin to how humans sketch or visualize future actions. We introduce a novel reinforcement learning framework, Visual Planning via Reinforcement Learning (VPRL), empowered by GRPO for post-training large vision models, leading to substantial improvements in planning in a selection of representative visual navigation tasks, FrozenLake, Maze, and MiniBehavior. Our visual planning paradigm outperforms all other planning variants that conduct reasoning in the text-only space. Our results establish Visual Planning as a viable and promising supplement to language-based reasoning, opening new avenues for tasks that benefit from intuitive, image-based inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。