arXiv:2507.04448cs.LGcond-mat.dis-nn2025-07

理论上揭示无限宽网络迁移学习何时能提升下游任务表现

Transfer Learning in Infinite Width Feature Learning Networks

  • 用梯度流分析无限宽网络的特征迁移机制
  • 预训练数据量与任务相关性影响下游性能,有明确可解释规律
  • 适合研究迁移学习理论或模型泛化能力的学者参考

我们建立了在梯度流下无限宽神经网络中迁移学习的理论,量化了在源任务上预训练是否能提升目标任务的泛化能力。分析两种情形:(i) 微调,即下游预测器在源任务生成的特征上训练;(ii) 共同丰富的设置,预训练和下游任务均可进入特征学习阶段,且下游模型初始化为预训练所得特征。此时,随机初始化网络在充分预训练后的统计量是依赖源数据与标签的自适应核。对 (i),分析不同预训练数据条件下读出层的表现;对 (ii),目标任务学习后的统计量仍为自适应核,包含源与目标任务的特征。我们在线性与多项式回归任务以及真实数据集上验证了该理论。结论可解释,依赖于两任务的数据量、任务间对齐程度及特征学习强度。

原文摘要 · Abstract (English)

We develop a theory of transfer learning in infinitely wide neural networks under gradient flow that quantifies when pretraining on a source task improves generalization on a target task. We analyze both (i) fine-tuning, when the downstream predictor is trained on top of source-induced features and (ii) a jointly rich setting, where both pretraining and downstream tasks can operate in a feature learning regime, but the downstream model is initialized with the features obtained after pre-training. In this setup, the summary statistics of randomly initialized networks after a rich pre-training are adaptive kernels which depend on both source data and labels. For (i), we analyze the performance of a readout for different pretraining data regimes. For (ii), the summary statistics after learning the target task are still adaptive kernels with features from both source and target tasks. We test our theory on linear and polynomial regression tasks as well as real datasets. Our theory allows interpretable conclusions on performance, which depend on the amount of data on both tasks, the alignment between tasks, and the feature learning strength.

迁移学习理论分析无限宽度泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。