arXiv:2606.23641cs.RO2026-06

通过保持模型平坦性,显著提升视觉语言动作模型对指令的遵循能力。

Flatness Preserves Instruction Following in Vision-Language-Action Models

论文配图:Flatness Preserves Instruction Following in Vision-Language-Action Models
图 1 · 摘自论文原文
  • finetune时引入尖锐度感知优化,使损失曲面更平坦
  • 在多个仿真与真实场景中指令遵循率提升超60%
  • 无需额外数据或修改结构,适合机器人控制研究者

视觉-语言-动作(VLA)模型可通过预训练的视觉-语言表征实现开放世界泛化,但下游在有限机器人数据上微调常导致表征退化,引发政策脆弱性,即忽略语言指令而依赖视觉捷径,这种现象称为指令盲视。我们假设:在数据有限时,标准微调对稀疏点施加梯度,形成高曲率极小值的尖锐损失曲面。为此,我们提出通过微调过程中的平坦性保持优化直接解决该问题,学习更平坦的损失曲面可增强权重空间扰动下的鲁棒性。具体而言,仅在微调中应用尖锐度感知最小化,即可在不增加数据、不修改架构或重训练的前提下,在多个仿真与真实世界基准上使指令遵循率提升超过60%。我们进一步分析了选择性尖锐度的影响,量化其效果,并证明该方法与现有引导技术具有互补性。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have the potential for open-world generalization by leveraging pretrained vision-language representations, yet downstream finetuning on limited robot data often degrades these representations, leading to brittle policies that ignore language instructions in favor of visual shortcuts, a failure mode we term instruction blindness. We hypothesize that standard finetuning with limited data applies gradients to a sparse set of points, which manifests as a sharp loss landscape with high-curvature minima. We propose to address this directly through flatness-preserving optimization while finetuning on the exact same data, where learning a flatter landscape results in a model more robust to perturbations in the weight space. Specifically, we demonstrate that simply applying sharpness-aware minimization during VLA finetuning significantly improves instruction following by over 60% across multiple simulation and real-world benchmarks without additional data, architectural modification, or retraining. We further analyze the effect of selective sharpness, quantify its effects, and show that our approach is complementary to existing guidance techniques. Project page can be found at https://haochenz11.github.io/papers/flatness-vla/.

VLA模型指令遵循平坦性优化机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。