通过分模块渐进放松约束,提升视觉语言动作模型的泛化能力
MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
- 按模块分阶段逐步放松预训练表示的约束,平衡稳定与适应性
- 在多个基准上实现最高30%的性能提升,覆盖仿真与真实机器人平台
- 无需新增参数或数据,可无缝接入现有视觉语言动作模型
视觉语言动作(VLA)模型继承自预训练视觉语言模型(VLM)的强大先验知识,但直接微调常破坏这些表示并损害泛化性能。现有方法如冻结模块或统一正则化,要么过度限制适应性,要么忽略各组件的不同作用。本文提出MAPS(模块级接近度调度),首个针对VLA的鲁棒微调框架。通过系统分析,发现应按特定顺序逐步放松对预训练表示的接近约束,以平衡稳定性与灵活性。MAPS采用线性调度策略,使视觉编码器保持贴近预训练先验,而面向动作的语言层可更自由地适应。该方法不引入额外参数或数据,可无缝集成至现有VLA模型。在MiniVLA-VQ、MiniVLA-OFT、OpenVLA-OFT及复杂基准如SimplerEnv、CALVIN、LIBERO,以及真实世界中Franka Emika Panda平台上的评估均显示,MAPS持续提升分布内与分布外性能(最高+30%)。研究揭示,基于实证的对预训练VLM表示的接近性,是实现从VLM到VLA迁移时广泛泛化的简单而有效原则。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models inherit strong priors from pretrained Vision-Language Models (VLMs), but naive fine-tuning often disrupts these representations and harms generalization. Existing fixes -- freezing modules or applying uniform regularization -- either overconstrain adaptation or ignore the differing roles of VLA components. We present MAPS (Module-Wise Proximity Scheduling), the first robust fine-tuning framework for VLAs. Through systematic analysis, we uncover an empirical order in which proximity constraints should be relaxed to balance stability and flexibility. MAPS linearly schedules this relaxation, enabling visual encoders to stay close to their pretrained priors while action-oriented language layers adapt more freely. MAPS introduces no additional parameters or data, and can be seamlessly integrated into existing VLAs. Across MiniVLA-VQ, MiniVLA-OFT, OpenVLA-OFT, and challenging benchmarks such as SimplerEnv, CALVIN, LIBERO, as well as real-world evaluations on the Franka Emika Panda platform, MAPS consistently boosts both in-distribution and out-of-distribution performance (up to +30%). Our findings highlight empirically guided proximity to pretrained VLMs as a simple yet powerful principle for preserving broad generalization in VLM-to-VLA transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。