arXiv:2601.23031stat.MLcs.LG2026-01

提出迭代经验风险最小化理论,揭示主动学习中标签预算分配的权衡

Asymptotic Theory of Iterated Empirical Risk Minimization, with Applications to Active Learning

  • 两阶段迭代训练:第一阶段预测结果作为第二阶段损失函数输入
  • 高维下精确预测测试误差,发现数据选择引发双下降现象
  • 适用于主动学习场景,无需先验假设或样本分割

我们研究一类迭代经验风险最小化(ERM)方法,其中在相同数据集上执行两次连续的ERM,且第一阶段估计器的预测结果作为第二阶段损失函数的输入。该设定自然出现在主动学习和重加权方案中,引入了样本间的复杂统计依赖关系,从根本上区别于经典单阶段ERM分析。针对高维情形下样本量与环境维度按比例增长的情况,我们在高斯混合数据上对线性模型、广泛凸损失函数下的测试误差给出了精确渐近刻画。即使存在数据重复使用和预测依赖损失,本理论仍能提供第二阶段估计器性能的完整渐近预测。我们将该理论应用于经典的池基主动学习问题,去除了先前工作中依赖的先验信息和样本分割假设。我们发现标签预算在各阶段间分配存在根本性权衡,并揭示了仅由数据选择驱动的双下降行为,而非模型规模或样本数量。

原文摘要 · Abstract (English)

We study a class of iterated empirical risk minimization (ERM) procedures in which two successive ERMs are performed on the same dataset, and the predictions of the first estimator enter as an argument in the loss function of the second. This setting, which arises naturally in active learning and reweighting schemes, introduces intricate statistical dependencies across samples and fundamentally distinguishes the problem from classical single-stage ERM analyses. For linear models trained with a broad class of convex losses on Gaussian mixture data, we derive a sharp asymptotic characterization of the test error in the high-dimensional regime where the sample size and ambient dimension scale proportionally. Our results provide explicit, fully asymptotic predictions for the performance of the second-stage estimator despite the reuse of data and the presence of prediction-dependent losses. We apply this theory to revisit a well-studied pool-based active learning problem, removing oracle and sample-splitting assumptions made in prior work. We uncover a fundamental tradeoff in how the labeling budget should be allocated across stages, and demonstrate a double-descent behavior of the test error driven purely by data selection, rather than model size or sample count.

主动学习双下降高维统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。