arXiv:2502.07016stat.MLcs.LG2025-02被引 1

为数据挖掘评估指标提供快速置信区间,提升结果可信度

Confidence Intervals for Evaluation of Data Mining

  • 基于渐近正态性,无需重抽样即可快速计算评估指标置信区间
  • 提出'模糊修正'方法,显著改善小样本下置信区间的覆盖概率
  • 支持多个分类器、多个指标同时推断,便于系统性比较

在数据挖掘中,二元预测规则用于预测二元结果时,常使用分类准确率、精确率、召回率、F值和杰卡德指数等性能度量进行评估与比较。这些度量通常仅在有限数据集上近似估计,可能导致统计不显著的结论。为量化这种不确定性,需为估计的性能度量提供置信区间。本文研究数据挖掘中通用性能度量的统计推断,涵盖个体与联合置信区间。置信区间基于渐近正态近似,可快速计算,无需自助法重抽样。我们分析了这些区间的有限样本覆盖概率,并提出一种'模糊修正'对方差进行调整,以改进小样本表现。该修正将二项比例的加四法推广至数据挖掘中的通用性能度量。本框架可同时对多个分类规则的多个性能度量进行推断,支持系统性比较。

原文摘要 · Abstract (English)

In data mining, when binary prediction rules are used to predict a binary outcome, many performance measures are used in a vast array of literature for the purposes of evaluation and comparison. Some examples include classification accuracy, precision, recall, F measures, and Jaccard index. Typically, these performance measures are only approximately estimated from a finite dataset, which may lead to findings that are not statistically significant. In order to properly quantify such statistical uncertainty, it is important to provide confidence intervals associated with these estimated performance measures. We consider statistical inference about general performance measures used in data mining, with both individual and joint confidence intervals. These confidence intervals are based on asymptotic normal approximations and can be computed fast, without needs to do bootstrap resampling. We study the finite sample coverage probabilities for these confidence intervals and also propose a `blurring correction' on the variance to improve the finite sample performance. This 'blurring correction' generalizes the plus-four method from binomial proportion to general performance measures used in data mining. Our framework allows multiple performance measures of multiple classification rules to be inferred simultaneously for comparisons.

性能评估置信区间数据挖掘统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。