提出DeLock方法,解决小数据微调后视觉语言动作模型失去指令响应能力的问题。
Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training

- 通过保留视觉锚定和测试时对比提示引导,防止模型过度适应训练数据。
- 在8个仿真与真实场景中表现优于基线,达到更高质量演示数据训练的效果。
- 无需额外监督信号或数据增强,适合资源有限的部署场景。
你是否曾在用少量示范数据对通用视觉-语言-动作(VLA)策略进行微调后,发现其不再响应新指令,仅能执行训练中见过的行为?我们识别出这一现象为‘锁死’:在低数据监督微调(SFT)后,策略过度特化于训练数据,无法泛化到新指令,表现为概念锁死(固定于训练对象/属性)和空间锁死(固定于训练空间目标)。现有许多修复方法依赖额外监督信号(如来自基础模型或辅助目标)或数据增强来恢复泛化能力。本文表明,策略内部预训练知识已足够:DeLock通过在微调过程中保持视觉锚定,并在测试时应用对比提示引导,根据新指令调整策略的去噪动态,从而缓解锁死。在8个仿真与真实世界评估中,DeLock持续优于强基线,表现达到甚至超过使用大量精调示范数据训练的最先进通用策略。
原文摘要 · Abstract (English)
Have you ever post-trained a generalist vision-language-action (VLA) policy on a small demonstration dataset, only to find that it stops responding to new instructions and is limited to behaviors observed during post-training? We identify this phenomenon as lock-in: after low-data, supervised fine-tuning (SFT), the policy becomes overly specialized to the post-training data and fails to generalize to novel instructions, manifesting as concept lock-in (fixation on training objects/attributes) and spatial lock-in (fixation on training spatial targets). Many existing remedies introduce additional supervision signals, such as those derived from foundation models or auxiliary objectives, or rely on augmented datasets to recover generalization. In this paper, we show that the policy's internal pre-trained knowledge is sufficient: DeLock mitigates lock-in by preserving visual grounding during post-training and applying test-time contrastive prompt guidance to steer the policy's denoising dynamics according to novel instructions. Across eight simulation and real-world evaluations, DeLock consistently outperforms strong baselines and matches or exceeds the performance of a state-of-the-art generalist policy post-trained with substantially more curated demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。