解析迁移学习在两类线性模型中的误差边界,给出辅助任务何时有效及如何加权的理论依据。
Expectation Error Bounds for Transfer Learning in Linear Regression and Linear Neural Networks
- 通过偏差-方差分解推导出线性回归的精确误差公式
- 在共享表示的线性神经网络中首次获得非渐近期望误差上界
- 提出可计算的最优任务权重与数据驱动的权重设计方法
在迁移学习中,学习者利用辅助数据提升主任务的泛化能力。然而,对辅助数据何时以及如何帮助仍缺乏精确的理论理解。本文针对两类典型线性场景——普通最小二乘回归和参数量不足的线性神经网络——提供了新见解。对于线性回归,我们推导出具有偏差-方差分解的期望泛化误差的精确闭式表达式,得出辅助任务改善主任务泛化的充要条件,并通过可解优化问题得到全局最优任务权重,且其经验估计具有一致性保证。对于共享表示宽度为 $q \≤ K$($K$ 为辅助任务数)的线性神经网络,我们建立了非渐近的期望误差上界,首次给出该设置下有益迁移学习的非平凡充分条件,并提供有原则的任务权重设计方向。这一成果基于对随机矩阵的列级低秩扰动的新界,相比现有结果更保留了细粒度的列结构信息。实验在可控参数的合成数据上验证了结论。
原文摘要 · Abstract (English)
In transfer learning, the learner leverages auxiliary data to improve generalization on a main task. However, the precise theoretical understanding of when and how auxiliary data help remains incomplete. We provide new insights on this issue in two canonical linear settings: ordinary least squares regression and under-parameterized linear neural networks. For linear regression, we derive exact closed-form expressions for the expected generalization error with bias-variance decomposition, yielding necessary and sufficient conditions for auxiliary tasks to improve generalization on the main task. We also derive globally optimal task weights as outputs of solvable optimization programs, with consistency guarantees for empirical estimates. For linear neural networks with shared representations of width $q \leq K$, where $K$ is the number of auxiliary tasks, we derive a non-asymptotic expectation bound on the generalization error, yielding the first non-vacuous sufficient condition for beneficial auxiliary learning in this setting, as well as principled directions for task weight curation. We achieve this by proving a new column-wise low-rank perturbation bound for random matrices, which improves upon existing bounds by preserving fine-grained column structures. Our results are verified on synthetic data simulated with controlled parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。