特征匹配比任务相似性更重要,预训练模型的特征空间重叠决定迁移效果。
Features are fate: a theory of transfer learning in high-dimensional regression
- 以特征空间重叠为核心,建立迁移学习理论框架。
- 在数据少时,特征重叠强则迁移优于从零训练。
- 结论适用于线性与非线性网络,对模型设计有指导意义。
随着大规模预训练神经网络的兴起,将此类‘基础’模型适配到数据有限的下游任务变得至关重要。当目标任务与源任务相近时,微调、偏好优化和迁移学习均取得成功,但对‘任务相似性’的精确理论理解仍不充分。传统观点认为,源与目标分布间的简单相似性度量(如ϕ-散度或积分概率度量)可直接预测迁移成功率,但我们证明这通常不成立。本文转而采用特征中心视角,建立一系列理论结果,表明当目标任务能被预训练模型的特征空间良好表征时,迁移学习性能超越从零训练。通过深度线性网络这一最小迁移学习模型,我们可解析刻画转移能力相图,其依赖于目标数据集规模与特征空间重叠程度。严格证明:当源-目标特征空间重叠足够强时,线性迁移与微调均提升性能,尤其在低数据条件下。这些结果基于对深度线性网络中特征学习动态的新兴理解,并数值验证了线性情况下的严谨结论同样适用于非线性网络。
原文摘要 · Abstract (English)
With the emergence of large-scale pre-trained neural networks, methods to adapt such "foundation" models to data-limited downstream tasks have become a necessity. Fine-tuning, preference optimization, and transfer learning have all been successfully employed for these purposes when the target task closely resembles the source task, but a precise theoretical understanding of "task similarity" is still lacking. While conventional wisdom suggests that simple measures of similarity between source and target distributions, such as $ϕ$-divergences or integral probability metrics, can directly predict the success of transfer, we prove the surprising fact that, in general, this is not the case. We adopt, instead, a feature-centric viewpoint on transfer learning and establish a number of theoretical results that demonstrate that when the target task is well represented by the feature space of the pre-trained model, transfer learning outperforms training from scratch. We study deep linear networks as a minimal model of transfer learning in which we can analytically characterize the transferability phase diagram as a function of the target dataset size and the feature space overlap. For this model, we establish rigorously that when the feature space overlap between the source and target tasks is sufficiently strong, both linear transfer and fine-tuning improve performance, especially in the low data limit. These results build on an emerging understanding of feature learning dynamics in deep linear networks, and we demonstrate numerically that the rigorous results we derive for the linear case also apply to nonlinear networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。