通过优化语言指令空间,让冻结的视觉语言动作模型更听话。
VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models

- 用强化学习优化语言指令,生成更有效的机器人操作命令。
- 在RL4VLA和VL-Think上,成功率提升23%~40%。
- 适合想提升机器人指令理解力的研究者与工程师。
视觉-语言-动作(VLA)模型通常被当作端到端的动作策略,以自然语言任务描述为条件。然而实践中,其行为对指令表述极为敏感,表明语言不仅是任务标签,更是可优化的输入变量。本文提出一种语言条件空间优化方法:不更新动作权重,而是优化语言空间,将人类指令转化为包含物体外观、空间关系和目标定位线索的短命令。该语言条件空间策略初始化于失败案例生成的命令先验,并在稀疏任务完成奖励下通过强化学习优化,下游VLA模型始终保持冻结。实验在RL4VLA和VL-Think数据集上验证,该方法显著提升了对指令敏感、符号化及多物体操作任务的成功率,证明语言可作为机器人基础模型的可优化变量。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. In practice, however, their behavior often depends sharply on how the instruction is phrased, suggesting that language is not merely a task label but an optimizable conditioning input. We study whether frozen VLA policies can be improved by optimizing language space rather than updating action weights. Our method introduces a language-conditioning space policy that translates a human instruction into a short VLA-grounded command using object appearance, spatial relations, and target-grounding cues. The language-conditioning space policy is initialized with a failure-derived command-space prior and optimized with reinforcement learning from sparse task-completion rewards, while the downstream VLA remains fully frozen. This yields language-conditioning space optimization: RL discovers which VLA-grounded commands best elicit successful behavior from the frozen action policy. Experiments on RL4VLA and VL-Think show that language-conditioning space optimization improves success on instruction-sensitive, symbolic, and multi-object manipulation tasks, demonstrating that language can serve as an optimizable variable for a robot foundation models. Website: https://tttonyalpha.github.io/vla_grounder
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。