arXiv:2410.22264cs.LG2024-10NeurIPS被引 1

用低秩适配提升模型元学习能力,理论证明比传统微调更优。

Provable Meta-Learning with Low-Rank Adaptations

  • 基于低秩适配(LoRA)设计可证明的元学习框架
  • 理论证明标准微调在适配新任务时必然次优
  • 实测在视觉与语言任务中显著优于传统方法

基础模型的强大之处在于其能学习到高度表达性的特征表示,并可适配于多种下游任务。然而,这些预训练模型需经过额外训练阶段才能有效应用。在多任务场景中,已有研究通过实验表明,特定的元学习方法结合参数高效微调(PEFT)可优于标准重训练,但其优势机制尚不明确。本文提出一种通用的基于PEFT的元学习框架,旨在学习一个易于适应未见任务的模型。针对使用LoRA的线性模型,我们证明标准重训练在寻找可适配参数方面是严格次优的,并为所提方法提供了严格的性能保证。通过合成数据及真实视觉和语言任务的实验验证了这些理论洞见。结果表明,相较于传统方法,采用简单实现的元学习方案在重训练过程中展现出显著性能提升。

原文摘要 · Abstract (English)

The power of foundation models (FMs) lies in their capacity to learn highly expressive representations that can be adapted to a broad spectrum of tasks. However, these pretrained models require additional training stages to become effective for downstream applications. In the multi-task setting, prior works have shown empirically that specific meta-learning approaches for preparing a model for future adaptation through parameter-efficient fine-tuning (PEFT) can outperform standard retraining methods, but the mechanism of the benefits of meta-learning has been largely unexplored. We introduce a framework for generic PEFT-based meta-learning to learn a model that can easily adapt to unseen tasks. For linear models using LoRA, we show that standard retraining is provably suboptimal for finding an adaptable set of parameters and provide strict performance guarantees for our proposed method. We verify these theoretical insights through experiments on synthetic data as well as real-data vision and language tasks. We observe significant performance benefits using a simple implementation of our proposed meta-learning scheme during retraining relative to the conventional approach.

元学习低秩适配参数高效微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。