arXiv:2505.03992cs.LG2025-05被引 1

小样本下分类指标因组合数学产生偏差,影响公平性评估

Algorithmic Accountability in Small Data: Sample-Size-Induced Bias Within Classification Metrics

  • 发现分类指标在小样本时受组合数学影响产生系统偏差
  • 揭示多个常用指标在群体规模不同时评估结果不可靠
  • 提出无需模型的偏差检测与修正方法,适合社会应用研究者

机器学习模型评估不仅关乎技术精度,也涉及潜在社会影响。尽管小样本偏差已广为人知,本文揭示了分类指标中由组合数学引发的样本量偏差具有重要影响。这一发现挑战了现有指标在高分辨率评估偏见上的有效性,尤其在群体规模差异大的社会应用场景中。我们分析了多个常用指标中的偏差现象,并提出一种模型无关的评估与修正技术。此外,还研究了指标计算中未定义情况的数量,若处理不当会导致误导性结论。本工作揭示了标准评估实践中被忽视的组合与概率问题,推动更公平可信的分类方法发展。

原文摘要 · Abstract (English)

Evaluating machine learning models is crucial not only for determining their technical accuracy but also for assessing their potential societal implications. While the potential for low-sample-size bias in algorithms is well known, we demonstrate the significance of sample-size bias induced by combinatorics in classification metrics. This revelation challenges the efficacy of these metrics in assessing bias with high resolution, especially when comparing groups of disparate sizes, which frequently arise in social applications. We provide analyses of the bias that appears in several commonly applied metrics and propose a model-agnostic assessment and correction technique. Additionally, we analyze counts of undefined cases in metric calculations, which can lead to misleading evaluations if improperly handled. This work illuminates the previously unrecognized challenge of combinatorics and probability in standard evaluation practices and thereby advances approaches for performing fair and trustworthy classification methods.

模型评估公平性小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。