arXiv:2508.04409stat.MLcs.LG2025-08被引 1

CV评估模型比较时可能失效,即使单个模型稳定。

The Relative Instability of Model Comparison with Cross-validation

  • 证明简单稳定模型间比较仍不稳
  • Lasso与软阈值法在理想条件下也导致无效CV
  • 提醒使用CV前必须验证相对稳定性

交叉验证(CV)被广泛用于模型比较,可提供渐近精确的检验和置信区间,但前提是模型比较本身相对稳定。本文出人意料地证明,即使单个模型表现稳定,其间的比较仍可能相对不稳定,从而质疑了CV推断的有效性。具体而言,我们证明了套索回归(Lasso)及其密切相关的软阈值方法,在最有利的学习环境下,即使两个模型均单独稳定,也会产生相对不稳定的比较结果,导致交叉验证推断失效。该发现强调了在部署交叉验证进行模型比较前,必须先验证其相对稳定性。

原文摘要 · Abstract (English)

Cross-validation (CV) is known to provide asymptotically exact tests and confidence intervals for model improvement but only when the model comparison is relatively stable. Surprisingly, we prove that even simple, individually stable models can generate relatively unstable comparisons, calling into question the validity of CV inference. Specifically, we show that the Lasso and its close cousin, soft-thresholding, generate relatively unstable comparisons and invalid CV inferences, even in the most favorable of learning settings and when both models are individually stable. These findings highlight the importance of verifying relative stability before deploying CV for model comparison.

交叉验证模型比较稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。