arXiv:2606.27295cs.RO2026-06被引 6

让机器人不看也能听懂指令执行动作,提升泛化能力

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

论文配图:LA4VLA: Learning to Act without Seeing via Language-Action Pretraining
图 1 · 摘自论文原文
  • 用语言描述动作片段替代视觉信号,构建无视觉依赖的预训练框架
  • 在仿真和真实任务中,成功率比基线最高提升45个百分点
  • 适合需要强语言理解与跨场景泛化的机器人控制研究

视觉-语言-动作(VLA)模型通常通过联合映射视觉观察和语言指令来生成动作进行预训练。然而,密集的视觉-动作监督会压倒相对稀疏的语言-动作信号,导致策略依赖视觉捷径而非语言指导,对视觉变化敏感。为此,我们提出LA4VLA,一种无需视觉观测即可学习语言条件动作先验的预训练框架。该框架将专家示范轨迹分解为原子动作片段,并配以低级动作描述,构建了33K条完全基于现有示范、无需额外数据采集的语言-动作(LA)数据集。我们进一步开发了10亿参数的轻量级模型LA4VLA-1B,探索三种语言-动作监督融入VLA学习的范式:仅语言预训练、语言到视觉-语言逐步预训练,以及混合语言-视觉-语言预训练。在仿真和真实世界任务中,语言预训练策略始终优于对应视觉-语言预训练基线,而混合预训练带来进一步提升:在仿真和真实任务中,平均成功率分别比无预训练基线提高17.8和45.0个百分点。结果表明,LA4VLA是一种有效且互补的预训练策略,有助于构建更强更鲁棒的VLA策略。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visual-action supervision can dominate the comparatively sparse language-action signal. As a result, policies may rely on visual shortcuts rather than learn how language conditions action execution, making them sensitive to visual variations. To address this limitation, we propose LA4VLA, a language-action pretraining framework that enables policies to acquire language-conditioned action priors without visual observations. These priors capture reusable manipulation skills shared across tasks and scenes, reducing reliance on scene-specific visual cues. Specifically, LA4VLA decomposes expert demonstration trajectories into atomic action segments and pairs each segment with a corresponding low-level action description. This yields LA-33K, a dataset of 33K Language-Action (LA) episodes derived entirely from existing demonstrations without additional robot data collection. We further develop LA4VLA-1B, a lightweight 1B-parameter VLA model, and investigate three paradigms for incorporating language-action supervision into VLA learning: LA-only pretraining, sequential LA-to-VLA pretraining, and mixed LA-VLA pretraining. Across simulation and real-world tasks, LA-pretrained policies consistently outperform matched VLA-pretrained counterparts, while combining LA and VLA supervision leads to further gains. In particular, mixed LA-VLA pretraining improves the average success rate of LA4VLA-1B over the no-pretraining baseline by up to 17.8 and 45.0 percentage points in simulation and real-world tasks, respectively. These results establish LA4VLA as an effective and complementary pretraining strategy for building stronger and more robust VLA policies.

语言动作机器人控制预训练泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。