arXiv:2410.02164cs.LGstat.ML2024-10NeurIPS被引 8

线性模型迁移学习的通用性分析,揭示微调为何更优。

Universality in Transfer Learning for Linear Models

  • 基于大模型渐近分析,仅依赖目标分布的一阶二阶统计量。
  • 在回归与分类任务中,微调模型泛化误差低于预训练模型。
  • 结果不依赖高斯假设,适用于多种优化方法和任务类型。

我们研究了线性模型在回归和二分类任务中的迁移学习与微调问题。重点关注使用小规模目标分布数据集,通过随机梯度下降(SGD)对预训练权重初始化的线性模型进行微调。在大模型渐近极限下,我们提供了严格且精确的分析,建立了预训练模型与微调模型在泛化误差(回归)和分类误差(二分类)之间的关系。特别地,给出了微调模型优于预训练模型的条件。本文的重要特点是所有结果具有'通用性',仅依赖于目标分布的一阶和二阶统计量,因此可超越文献中常见的高斯假设。此外,我们的通用性结论也扩展至使用岭回归训练的分类任务的测试误差。

原文摘要 · Abstract (English)

We study the problem of transfer learning and fine-tuning in linear models for both regression and binary classification. In particular, we consider the use of stochastic gradient descent (SGD) on a linear model initialized with pretrained weights and using a small training data set from the target distribution. In the asymptotic regime of large models, we provide an exact and rigorous analysis and relate the generalization errors (in regression) and classification errors (in binary classification) for the pretrained and fine-tuned models. In particular, we give conditions under which the fine-tuned model outperforms the pretrained one. An important aspect of our work is that all the results are "universal", in the sense that they depend only on the first and second order statistics of the target distribution. They thus extend well beyond the standard Gaussian assumptions commonly made in the literature. Furthermore, our universality results extend beyond standard SGD training to the test error of a classification task trained using a ridge regression.

迁移学习线性模型通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。