研究数据增强对随机特征回归泛化误差的影响,给出精确的误差预测公式。
Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation

- 基于数据分布和增强统计量,推导测试误差的闭式表达
- 在特征映射错误设定下仍保持高精度预测,误差与真实数据一致
- 适用于仅训练输出层的模型,适合机器学习理论研究者
本文旨在分析在比例样本-特征情形(即协变量数量与样本量同比增加)下,数据增强对监督回归方法所引入的正则化效应。我们仅基于真实数据的总体统计量,以及增强方案的一阶和二阶统计量,给出了测试误差(均方误差度量)的紧致刻画。该结果在特征映射不正确的情况下依然成立,并适用于仅训练最后一层输出层、其余网络部分被冻结或随机初始化的任意网络结构。针对高斯数据情形,我们具体给出了结果,并证明在该设定下我们的渐近刻画是紧的。
原文摘要 · Abstract (English)
This paper aims at analyzing the regularization effect that data augmentation induces on supervised regression methods in the proportional regime, where the number of covariates grows proportionally to the number of samples. We provide a tight characterization of the test error, measured in mean squared error, in terms only of the population quantities of the true data, as well as first and second order statistics of the augmentation scheme. Our results are valid under misspecified feature maps, and for any network architecture where only the last readout layer is trained, and the rest of the network is either frozen or randomly initialized. We specify our results in the case of Gaussian data, and show that our asymptotic characterization is tight in this setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。