arXiv:2604.17694stat.MEcs.LG2026-04被引 1

通过袋装法提升机器学习估计的随机种子稳定性,增强可复现性。

Improving reproducibility by controlling random seed stability in machine learning based estimation via bagging

  • 用袋装法控制随机种子影响,确保结果稳定
  • 新方法在数值实验中实现目标稳定性,传统方法不行
  • 适合关注模型可复现性的研究者和实践者

机器学习算法的预测结果会因随机种子不同而波动,导致下游去偏估计器不稳定。本文通过集中条件形式化随机种子稳定性,并证明子袋装法对任意有界输出回归算法均能保证稳定性。提出一种新交叉拟合方法——自适应交叉袋装,同时消除噪声估计和样本分割中的种子依赖。数值实验表明,该方法能达到预期稳定性水平,而其他方法不能。相比标准做法,本方法仅带来小幅计算开销,而其他方法则成本巨大。

原文摘要 · Abstract (English)

Predictions from machine learning algorithms can vary across random seeds, inducing instability in downstream debiased machine learning estimators. We formalize random seed stability via a concentration condition and prove that subbagging guarantees stability for any bounded-outcome regression algorithm. We introduce a new cross-fitting procedure, adaptive cross-bagging, which simultaneously eliminates seed dependence from both nuisance estimation and sample splitting in debiased machine learning. Numerical experiments confirm that the method achieves the targeted level of stability whereas alternatives do not. Our method incurs a small computational penalty relative to standard practice whereas alternative methods incur large penalties.

可复现性机器学习去偏估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。