arXiv:2510.07649stat.MLcs.LG2025-10

提出新方法,更准确评估特定模型的真实预测性能。

A Honest Cross-Validation Estimator for Prediction Performance

  • 基于随机效应模型,用其他数据分拆改进单一测试集估计。
  • 模拟与真实数据均显示优于传统交叉验证和单次分割法。
  • 适合关注模型实际表现而非平均性能的研究者使用。

交叉验证是评估预测模型性能的标准工具。常用方法是多次随机划分数据,用训练集训练模型,测试集评估性能,并对不同划分结果取均值。一个广为人知的批评是:该过程并未直接估计最终推荐模型的实际性能。本文提出一种新方法,用于估计在特定(随机)训练集上训练出的模型性能。朴素估计可通过将模型应用于独立测试集获得。令人意外的是,在随机效应模型框架下,可利用其他随机划分产生的交叉验证结果来改进这一朴素估计。我们发展了两种估计器——层次贝叶斯估计器和经验贝叶斯估计器——其性能与或优于传统交叉验证估计器及单次分割估计器。模拟实验与真实数据案例均验证了所提方法的优越性。

原文摘要 · Abstract (English)

Cross-validation is a standard tool for obtaining a honest assessment of the performance of a prediction model. The commonly used version repeatedly splits data, trains the prediction model on the training set, evaluates the model performance on the test set, and averages the model performance across different data splits. A well-known criticism is that such cross-validation procedure does not directly estimate the performance of the particular model recommended for future use. In this paper, we propose a new method to estimate the performance of a model trained on a specific (random) training set. A naive estimator can be obtained by applying the model to a disjoint testing set. Surprisingly, cross-validation estimators computed from other random splits can be used to improve this naive estimator within a random-effects model framework. We develop two estimators -- a hierarchical Bayesian estimator and an empirical Bayes estimator -- that perform similarly to or better than both the conventional cross-validation estimator and the naive single-split estimator. Simulations and a real-data example demonstrate the superior performance of the proposed method.

交叉验证模型评估贝叶斯方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。