arXiv:2410.08455cs.LGcs.AI2024-10被引 5

预训练通过保留关键知识,让模型更快更准地学会下游任务。

Why pre-training is beneficial for downstream classification tasks?

  • 从博弈论视角量化分析预训练知识的迁移机制。
  • 仅少量预训练知识被保留,但极难从零学习获得。
  • 帮助模型直接快速掌握目标任务,加速收敛。

预训练在提升下游分类任务准确率和加速收敛方面表现出显著优势,但其内在原因仍不明确。为此,我们提出一种新颖的博弈论视角,定量且显式地解释预训练对下游任务的影响,同时为深度神经网络的学习行为提供了新见解。具体而言,我们提取并量化了预训练模型所编码的知识,并追踪其在微调过程中的变化。令人惊讶的是,仅有少量预训练知识被保留在下游任务推理中,但这些知识对于从零训练的模型来说极为难以习得。因此,借助这一独特且有用的知识,微调模型通常表现优于从头训练的模型。此外,我们发现预训练能引导微调模型更直接、更快速地学习目标任务知识,这解释了微调模型收敛更快的原因。

原文摘要 · Abstract (English)

Pre-training has exhibited notable benefits to downstream tasks by boosting accuracy and speeding up convergence, but the exact reasons for these benefits still remain unclear. To this end, we propose to quantitatively and explicitly explain effects of pre-training on the downstream task from a novel game-theoretic view, which also sheds new light into the learning behavior of deep neural networks (DNNs). Specifically, we extract and quantify the knowledge encoded by the pre-trained model, and further track the changes of such knowledge during the fine-tuning process. Interestingly, we discover that only a small amount of pre-trained model's knowledge is preserved for the inference of downstream tasks. However, such preserved knowledge is very challenging for a model training from scratch to learn. Thus, with the help of this exclusively learned and useful knowledge, the model fine-tuned from pre-training usually achieves better performance than the model training from scratch. Besides, we discover that pre-training can guide the fine-tuned model to learn target knowledge for the downstream task more directly and quickly, which accounts for the faster convergence of the fine-tuned model.

预训练知识迁移微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。