让视觉语言动作模型更听话:自动优化指令,避免乱改导致失败
Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering

- 交互式搜索有效指令,生成可复用的反馈策略
- 仿真中提升24.7%,硬件实验提升65.0%
- 保证不伤害模型表现,适合真实机器人部署
视觉-语言-动作(VLA)模型为机器人控制提供了自然语言接口,但语言到行为的映射常不稳定且不可预测:语义相近的指令可能引发截然不同的行为,部分能力仅靠提示无法激发。为此,我们提出一种框架,通过交互式搜索提升闭环任务性能的语言序列,将其提炼为测试时的语言反馈策略(LFP),并训练一个改进头预测语言引导是否有效。我们对改进头进行置信区间校准,防止在分布外场景下因语言干预导致性能下降。关键优势在于无需访问原始训练数据或微调底层模型,适用于任意预训练冻结的VLA。在已知环境中,校准后的LFP使模拟环境性能提升24.7%,硬件实验提升65.0%;在视觉与语义扰动下,该方法具备强安全性保障,能生成开环提示无法实现的恢复行为。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone. As a result, both human instructions and zero-shot language models can fail to reliably steer VLAs toward successful task execution. In this work, we propose a framework that interactively searches for language sequences that improve closed-loop VLA task performance, distills these sequences into a test-time language feedback policy (LFP), and learns an improvement head that predicts when language steering will improve performance. We conformalize this improvement head to prevent harmful steering interventions, where the LFP decreases task performance relative to the original instruction on out-of-distribution scenarios. Crucially, our approach operates on arbitrary frozen pre-trained VLAs, requiring neither access to the original training distribution nor fine-tuning of the underlying model. On seen environments, our conformalized LFP improves base VLA performance by 24.7% in simulation and 65.0% in hardware. On visual and semantic perturbations, our conformalized LFP has strong harmlessness guarantees, and produces recovery behaviors not observed with open-loop prompting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。