arXiv:2602.10556cs.ROcs.AI2026-02被引 19

让机器人用自然语言理解动作,实现零样本跨平台通用控制

LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer

  • 用自然语言直接表示机器人动作,与视觉语言模型对齐
  • 在未见过的机器人上零样本成功率超50%,提升约2倍
  • 无需微调即可跨机器人部署,适合通用机器人研发

机器人领域的长期目标是构建可零样本部署于新机器人平台的通用策略,而无需针对每个平台进行适配。尽管已有大规模多平台预训练,现有视觉-语言-动作模型(VLAs)仍与其训练平台紧密绑定,通常需要昂贵的微调。本文提出语言-动作预训练(LAP),将底层机器人动作直接以自然语言表示,使动作监督与预训练视觉-语言模型的输入输出分布一致。LAP无需学习分词器、无需高成本标注,也无需特定平台的架构设计。基于LAP,我们构建了LAP-3B,据我们所知,这是首个在未经任何平台特化微调的情况下,实现显著零样本跨平台迁移的VLA。在多个新型机器人和操作任务中,LAP-3B平均零样本成功率超过50%,相较最强先前模型提升约2倍。我们还证明LAP支持高效适应和良好可扩展性,并通过共享语言-动作格式统一动作预测与视觉问答,实现协同训练增益。

原文摘要 · Abstract (English)

A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision-Language-Action models (VLAs) remain tightly coupled to their training embodiments and typically require costly fine-tuning. We introduce Language-Action Pre-training (LAP), a simple recipe that represents low-level robot actions directly in natural language, aligning action supervision with the pre-trained vision-language model's input-output distribution. LAP requires no learned tokenizer, no costly annotation, and no embodiment-specific architectural design. Based on LAP, we present LAP-3B, which to the best of our knowledge is the first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments without any embodiment-specific fine-tuning. Across multiple novel robots and manipulation tasks, LAP-3B attains over 50% average zero-shot success, delivering roughly a 2x improvement over the strongest prior VLAs. We further show that LAP enables efficient adaptation and favorable scaling, while unifying action prediction and VQA in a shared language-action format that yields additional gains through co-training.

机器人零样本迁移语言动作通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。