arXiv:2507.08224cs.RO2025-07EMNLP被引 4

让小模型自动生成可执行的机器人操作计划

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning

  • 小模型通过自我批判与修正生成更准确的执行计划
  • 3B到72B模型均提升性能,部分超越100倍大的模型
  • 无需外部监督,适合资源有限的机器人应用

大语言模型在机器人程序规划中展现潜力,但其人类中心推理常忽略执行所需的底层具体细节。视觉语言模型(VLMs)提供更贴近感知的规划路径,但现有方法要么依赖昂贵的大规模模型,要么局限于狭窄仿真环境。我们提出SelfReVision,一种轻量级、可扩展的自改进框架,用于视觉语言程序规划。SelfReVision使小型VLM能无需外部监督或教师模型,通过自批判、修订和验证自身计划,借鉴思维链提示与自指导范式。经此自蒸馏循环,模型生成更高质量、可执行的计划,既可用于推理,也可用于持续微调。使用3B至72B模型的实验表明,SelfReVision不仅显著提升弱基线VLM性能,还超越了规模达其100倍的模型,在下游具身任务中实现更好控制。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown promise in robotic procedural planning, yet their human-centric reasoning often omits the low-level, grounded details needed for robotic execution. Vision-language models (VLMs) offer a path toward more perceptually grounded plans, but current methods either rely on expensive, large-scale models or are constrained to narrow simulation settings. We introduce SelfReVision, a lightweight and scalable self-improvement framework for vision-language procedural planning. SelfReVision enables small VLMs to iteratively critique, revise, and verify their own plans-without external supervision or teacher models-drawing inspiration from chain-of-thought prompting and self-instruct paradigms. Through this self-distillation loop, models generate higher-quality, execution-ready plans that can be used both at inference and for continued fine-tuning. Using models varying from 3B to 72B, our results show that SelfReVision not only boosts performance over weak base VLMs but also outperforms models 100X the size, yielding improved control in downstream embodied tasks.

机器人规划视觉语言模型自蒸馏小模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。