arXiv:2602.03973cs.ROcs.CV2026-02被引 7

用视觉语言模型让机器人在测试时自动调整动作,应对环境变化。

VLS: Steering Pretrained Robot Policies via Vision-Language Models

  • 通过视觉语言模型生成可微奖励,动态引导生成动作
  • 在CALVIN上提升31%,LIBERO-PRO上提升13%
  • 无需重训,适合真实场景中突发环境变化

为何预训练的扩散或流匹配策略在靠近障碍物、支撑面偏移或轻微杂乱环境下会失效?这类失败通常并非因缺乏运动技能,而是模仿学习在训练-测试分布偏移下的局限性所致——动作生成与训练时的空间配置和任务设定高度耦合。重新训练或微调成本高且逻辑不符,因为所需行为已存在,仅需在测试时选择性适配。我们提出视觉语言引导(VLS)框架,实现冻结生成式机器人策略的推理时自适应。VLS将适应视为推理时的控制问题,通过视觉语言模型合成轨迹可微奖励函数,引导预训练扩散或流匹配策略的去噪过程,生成满足测试时空与任务需求的动作轨迹。在仿真与真实世界评估中,VLS持续优于已有引导方法,在CALVIN上提升31%,在LIBERO-PRO上提升13%。在Franka机器人上的真实部署也证明了其对测试时空间与语义偏移的鲁棒自适应能力。

原文摘要 · Abstract (English)

Why do pretrained diffusion or flow-matching policies fail when the same task is performed near an obstacle, on a shifted support surface, or amid mild clutter? Such failures rarely reflect missing motor skills; instead, they expose a limitation of imitation learning under train-test shifts, where action generation is tightly coupled to training-specific spatial configurations and task specifications. Retraining or fine-tuning to address these failures is costly and conceptually misaligned, as the required behaviors already exist but cannot be selectively adapted at test time. We propose Vision-Language Steering (VLS), a training-free framework for inference-time adaptation of frozen generative robot policies. VLS treats adaptation as an inference-time control problem, steering the sampling process of a pretrained diffusion or flow-matching policy in response to out-of-distribution observation-language inputs without modifying policy parameters. By leveraging vision-language models to synthesize trajectory-differentiable reward functions, VLS guides denoising toward action trajectories that satisfy test-time spatial and task requirements. Across simulation and real-world evaluations, VLS consistently outperforms prior steering methods, achieving a 31% improvement on CALVIN and a 13% gain on LIBERO-PRO. Real-world deployment on a Franka robot further demonstrates robust inference-time adaptation under test-time spatial and semantic shifts. Project page: https://vision-language-steering.github.io/webpage/

机器人视觉语言生成模型自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。