训练源模型最优反而不利于下游迁移,关键在任务对齐程度。
Source-Optimal Training is Transfer-Suboptimal
- 源模型最优正则化与迁移最优不一致,仅在极少数情况下相同。
- 任务对齐弱时需更强正则化,对齐强时反而需更弱正则化。
- 该现象在非线性模型中依然存在,适用于主流迁移学习场景。
我们证明,在下游迁移目标下,为源任务优化训练模型通常是次优的。研究了L2-SP岭回归中的源端优化问题,发现源最优正则化参数τ₀*与迁移最优正则化参数τ_S*在几乎所有情况下均不相等(除测度为零的情况外)。我们刻画了迁移最优的源惩罚τ₀*与任务对齐度的关系,并识别出一种依赖于对齐度的反转现象:当任务对齐度ρ在0到1之间(不完全对齐)时,更强的源正则化有利于迁移;而当ρ>1(超对齐)时,较弱的正则化更有利。此外,在各向同性设置下,迁移是否有益与目标样本量和噪声无关,仅取决于任务对齐度和源模型特性。我们在合成岭回归实验中验证了线性预测,同时在MNIST、CIFAR-10和20 Newsgroups数据集上的实验表明,这种源最优与迁移最优的不匹配在标准非线性迁移学习流程中仍然存在。
原文摘要 · Abstract (English)
We prove that training a source model optimally for its own task is generically suboptimal when the objective is downstream transfer. We study the source-side optimization problem in L2-SP ridge regression and show a fundamental mismatch between the source-optimal and transfer-optimal source regularization: outside of a measure-zero set, $τ_0^* \neq τ_S^*$. We characterize the transfer-optimal source penalty $τ_0^*$ as a function of task alignment and identify an alignment-dependent reversal: with imperfect alignment ($0<ρ<1$), transfer benefits from stronger source regularization, while in super-aligned regimes ($ρ>1$), transfer benefits from weaker regularization. Additionally, in isotropic settings, the decision of whether transfer helps is independent of the target sample size and noise, depending only on task alignment and source characteristics. We verify the linear predictions in a synthetic ridge regression experiment, and we present experiments on MNIST, CIFAR-10, and 20 Newsgroups as evidence that the source-optimal versus transfer-optimal mismatch persists in standard nonlinear transfer learning pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。