arXiv:2410.17128stat.MLcs.LG2024-10被引 3

用概率测度微分理论分析迁移学习泛化误差,揭示其有效性机制。

Understanding Transfer Learning via Mean-field Analysis

  • 基于概率测度空间上的微分计算,构建迁移学习理论框架。
  • 在均场神经网络下,证明了迁移学习可降低泛化误差并加速收敛。
  • 适合研究迁移学习理论机制的学者,尤其关注泛化性能分析者。

我们提出一种新框架,通过概率测度空间上的微分计算来探究迁移学习的泛化误差。具体考虑两种主要迁移学习场景:α-ERM 与带 KL 正则化的经验风险最小化微调,并建立了在损失函数与激活函数满足一定可积性与正则性条件下,泛化误差及总体风险收敛速率的通用条件。基于理论结果,我们在均场神经网络设置下,证明了在适当假设下,单隐层神经网络的迁移学习具有优势,能有效降低泛化误差并提升收敛速度。

原文摘要 · Abstract (English)

We propose a novel framework for exploring generalization errors of transfer learning through the lens of differential calculus on the space of probability measures. In particular, we consider two main transfer learning scenarios, $α$-ERM and fine-tuning with the KL-regularized empirical risk minimization and establish generic conditions under which the generalization error and the population risk convergence rates for these scenarios are studied. Based on our theoretical results, we show the benefits of transfer learning with a one-hidden-layer neural network in the mean-field regime under some suitable integrability and regularity assumptions on the loss and activation functions.

迁移学习均场理论泛化误差神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。