仅用一次数据集,高效评估大模型预测误差。
Interleaved Resampling and Refitting: Data and Compute-Efficient Evaluation of Black-Box Predictors
- 通过交替重采样与小规模重训练,实现低资源评估。
- 仅需少量小数据集重训,即可获得误差上界。
- 适合无法获取验证集的大模型评估场景。
我们研究在平方损失下大规模经验风险最小化方法的过拟合风险评估问题。基于野蛮重拟合与重采样的思想,假设仅能通过黑箱访问训练算法,提出一种高效估计过拟合风险的方法。该评估算法在计算和数据上均高效:仅需一个数据集,无需额外验证数据。计算上,只需对通过顺序重采样生成的多个小数据集进行多次小规模重训练,避免了以往方法所需的全规模重训练,适用于大规模训练模型。算法采用交错的序列重采样与重拟合结构:先通过随机残差对称化构造伪响应;每轮从生成的协变量-伪响应对中重采样两个子数据集;随后分别在两个小型人工数据集上独立重训练模型。在固定设计与随机设计下,理论证明了高概率的过拟合风险上界,表明在适当选择噪声尺度时,该算法可给出预测误差的上界。理论分析借助经验过程理论、调和分析、托普利茨算子理论及尖锐张量集中不等式。
原文摘要 · Abstract (English)
We study the problem of evaluating the excess risk of large-scale empirical risk minimization under the square loss. Leveraging the idea of wild refitting and resampling, we assume only black-box access to the training algorithm and develop an efficient procedure for estimating the excess risk. Our evaluation algorithm is both computationally and data efficient. In particular, it requires access to only a single dataset and does not rely on any additional validation data. Computationally, it only requires refitting the model on several much smaller datasets obtained through sequential resampling, in contrast to previous wild refitting methods that require full-scale retraining and might therefore be unsuitable for large-scale trained predictors. Our algorithm has an interleaved sequential resampling-and-refitting structure. We first construct pseudo-responses through a randomized residual symmetrization procedure. At each round, we thus resample two sub-datasets from the resulting covariate pseudo-response pairs. Finally, we retrain the model separately on these two small artificial datasets. We establish high probability excess risk guarantees under both fixed design and random design settings, showing that with a suitably chosen noise scale, our interleaved resampling and refitting algorithm yields an upper bound on the prediction error. Our theoretical analysis draws on tools from empirical process theory, harmonic analysis, Toeplitz operator theory, and sharp tensor concentration inequalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。