arXiv:2603.16044cs.AI2026-03

用合成指令提升机器人模型对新环境的语言理解能力

Enhancing Linguistic Generalization of VLA: Fine-Tuning OpenVLA via Synthetic Instruction Augmentation

  • 用大模型生成结构多样的等义指令扩充数据集
  • LoRA微调后模型在新环境零样本表现显著提升
  • 适合关注机器人语言泛化与高效微调的研究者

具身智能中的泛化仍是核心挑战,机器人需适应多样环境。尽管OpenVLA通过大规模预训练成为视觉-语言-动作模型的最新标杆,但在完全陌生环境中其零样本性能仍受限。本文提出一种参数高效的微调策略,通过为Bridge Dataset V2合成通用指令集来增强OpenVLA的语言泛化能力。利用大语言模型(LLM)为现有轨迹生成大量语义等价但结构多样的指令。实验采用低秩适配(LoRA)在增广数据对上微调OpenVLA,使模型更有效地将复杂自然语言意图映射到机器人动作。结果表明,经LoRA增强的模型具备更强鲁棒性,说明丰富专业化数据集的语言空间对具身智能体至关重要。

原文摘要 · Abstract (English)

Generalization remains a core challenge in embodied AI, as robots must adapt to diverse environments. While OpenVLA represents the State-of-the-Art (SOTA) in Vision-Language-Action models by leveraging large-scale pre-training, its zero-shot performance can be limited when encountering completely new environments. This paper proposes a parameter-efficient fine-tuning strategy to enhance the linguistic generalization of OpenVLA by synthesizing a general instruction set for the Bridge Dataset V2. The paper leverages a Large Language Model (LLM) to generate a rich variety of semantically equivalent but structurally diverse commands for existing trajectories. In this experiment, Low-Rank Adaptation (LoRA) is implemented to fine-tune OpenVLA on augmented pairs, allowing the model to bridge the gap between complex natural language intent and robotic actions. Results demonstrate that the LoRA-enhanced model's robustness, suggesting that enriching the linguistic space of specialized datasets is crucial for embodied agents.

机器人语言泛化LoRA指令生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。