理论证明大模型能显著降低下游任务所需数据量。
Provable Target Sample Complexity Improvements as Pre-Trained Models Scale
- 基于参数高效微调思路构建新分析框架
- 证明模型越大,下游学习所需样本越少
- 为大模型缩放效应提供首个严格理论解释
预训练模型已成为高效构建下游任务模型的基石。已有实证研究揭示了缩放定律,表明更大的预训练模型可显著降低下游学习的样本复杂度。然而,现有理论研究尚无法解释该现象。本文提出一种受参数高效微调(如Adapter、LoRA、部分微调)启发的新理论框架caulking。分析结果证明:随着预训练模型规模增大,下游任务的样本复杂度可被严格证明地降低,从而为观测到的预训练模型规模与下游性能之间的缩放关系提供了首个理论支持,填补了现有研究空白。
原文摘要 · Abstract (English)
Pre-trained models have become indispensable for efficiently building models across a broad spectrum of downstream tasks. The advantages of pre-trained models have been highlighted by empirical studies on scaling laws, which demonstrate that larger pre-trained models can significantly reduce the sample complexity of downstream learning. However, existing theoretical investigations of pre-trained models lack the capability to explain this phenomenon. In this paper, we provide a theoretical investigation by introducing a novel framework, caulking, inspired by parameter-efficient fine-tuning (PEFT) methods such as adapter-based fine-tuning, low-rank adaptation, and partial fine-tuning. Our analysis establishes that improved pre-trained models provably decrease the sample complexity of downstream tasks, thereby offering theoretical justification for the empirically observed scaling laws relating pre-trained model size to downstream performance, a relationship not covered by existing results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。