arXiv:2502.16707cs.ROcs.AI2025-02被引 52

用反思机制提升视觉语言模型的长程机器人操作能力

Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation

  • 通过生成未来状态并反思错误,迭代优化动作决策
  • 在多阶段操作任务中显著优于现有商业VLM和MCTS方法
  • 适合需要复杂物理推理的长时序机器人控制场景

解决复杂的长时序机器人操作问题需要高级规划能力、对物理世界的理解以及动态选择合适运动技能的能力。虽然预训练于网络数据的视觉语言模型(VLMs)理论上可应对此类挑战,但当前版本缺乏精细的物理理解能力,且难以在长时序中有效应对误差累积问题。本文提出一种新型测试时计算框架,通过引入“反思”机制增强VLM的物理推理能力。该方法利用生成模型预测未来世界状态,据此指导动作选择,并通过批判性反思潜在次优性来持续改进推理过程。实验表明,该方法在多阶段操作任务中显著优于多个先进商业VLM及后训练方法(如蒙特卡洛树搜索)。视频演示见 https://reflect-vlm.github.io。

原文摘要 · Abstract (English)

Solving complex long-horizon robotic manipulation problems requires sophisticated high-level planning capabilities, the ability to reason about the physical world, and reactively choose appropriate motor skills. Vision-language models (VLMs) pretrained on Internet data could in principle offer a framework for tackling such problems. However, in their current form, VLMs lack both the nuanced understanding of intricate physics required for robotic manipulation and the ability to reason over long horizons to address error compounding issues. In this paper, we introduce a novel test-time computation framework that enhances VLMs' physical reasoning capabilities for multi-stage manipulation tasks. At its core, our approach iteratively improves a pretrained VLM with a "reflection" mechanism - it uses a generative model to imagine future world states, leverages these predictions to guide action selection, and critically reflects on potential suboptimalities to refine its reasoning. Experimental results demonstrate that our method significantly outperforms several state-of-the-art commercial VLMs as well as other post-training approaches such as Monte Carlo Tree Search (MCTS). Videos are available at https://reflect-vlm.github.io.

机器人操作视觉语言模型长程规划反思机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。