arXiv:2506.12409cs.CV2025-06AAAI被引 2

用零阶优化提升视觉语言模型持续学习效果

Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models

  • 从分支到层粒度探索零阶优化,寻找最优策略
  • 在四个基准上实现当前最佳性能,有效跳出局部最优
  • 提出模态感知的零阶方法,特别改善视觉模态优化

视觉-语言持续学习(VLCL)因强大泛化能力受到广泛关注,采用参数高效微调(PEFT)策略可显著降低资源消耗并保持竞争力。然而,主流的一阶(FO)优化易陷入次优局部极小值,尤其在PEFT受限的探索空间中表现不佳。本文首次系统研究将零阶(ZO)优化应用于基于PEFT的VLCL。我们发现直接全量使用ZO会导致优化过程不稳定,因此从模态分支级到细粒度层级逐步探索其适用性。理论分析揭示:在ZO优化过程中,视觉模态方差显著高于语言模态,为此提出一种模态感知的零阶策略,通过梯度符号归一化与限制视觉模态扰动来提升性能。得益于该策略,基于PEFT的VLCL具备更强逃离局部极小的能力。在四个基准上的广泛实验表明,本方法达到当前最优结果。

原文摘要 · Abstract (English)

Vision-Language Continual Learning (VLCL) has attracted significant research attention for its robust capabilities, and the adoption of Parameter-Efficient Fine-Tuning (PEFT) strategies is enabling these models to achieve competitive performance with substantially reduced resource consumption. However, dominated First-Order (FO) optimization is prone to trap models in suboptimal local minima, especially in limited exploration subspace within PEFT. To overcome this challenge, this paper pioneers a systematic exploration of adopting Zeroth-Order (ZO) optimization for PEFT-based VLCL. We first identify the incompatibility of naive full-ZO adoption in VLCL due to optimization process instability. We then investigate the application of ZO optimization from a modality branch-wise to a fine-grained layer-wise across various training units to identify an optimal strategy. Besides, a key theoretical insight reveals that vision modality exhibit higher variance than language counterparts in VLCL during the ZO optimization process, and we propose a modality-aware ZO strategy, which adopts gradient sign normalization in ZO and constrains vision modality perturbation to further improve performance. Benefiting from the adoption of ZO optimization, PEFT-based VLCL fulfills better ability to escape local minima during the optimization process, extensive experiments on four benchmarks demonstrate that our method achieves state-of-the-art results.

持续学习零阶优化视觉语言模型参数高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。