arXiv:2606.26422cs.AI2026-06被引 1

为小样本文本分类提供更准确的置信区间估算方法,提升大模型评估可靠性。

Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data

论文配图:Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data
图 1 · 摘自论文原文
  • 提出伪计数正则化自助法等改进方法,适配小样本与高精度场景。
  • 在嵌套数据下,需同时调整有效样本量和自由度以获准确区间。
  • 适用于社会科学研究中的文本分类验证,尤其关注样本设计阶段。

研究人员越来越多地使用文本分类模型(包括大语言模型)来测量自然语言中的概念,并以召回率、精确率等指标作为其有效性的证据。然而,这些指标是受抽样变异影响的点估计值,其不确定性度量却常被忽略或采用不恰当的方法,尤其是在标注数据集较小或性能较高时。本文评估了在社会科学文本分类典型条件下(小到中等样本量、罕见概念、文本嵌套于个体)性能指标的置信区间方法。模拟结果显示,传统的Wald区间和基本百分位自助法准确性最差,覆盖概率有时远低于名义95%水平。使用Agresti-Coull、Wilson、Clopper-Pearson及一种新型伪计数正则化自助法可显著提升准确性,其中后者特别适用于F1分数计算。当文本嵌套于个体时,必须同时调整有效样本量和适当的自由度才能获得准确的分析区间。在自助法中,分层自助法在个体产生中等数量文本时表现优于聚类自助法,但在个体文本数较少时过于保守。本文为领域提供合适的区间估计指导,旨在提升机器学习应用的透明性,并推动研究者在设计阶段更关注验证样本量。

原文摘要 · Abstract (English)

Researchers increasingly use text classification--supervised models or large language models--to measure constructs from natural language, providing metrics such as recall and precision as evidence of their validity. Yet, though these metrics are point estimates subject to sampling variation, measures of uncertainty are inconsistently reported alongside them. Further, when they are reported, they are often estimated with methods that are not appropriate when relevant labelled datasets are small or performance is high. To increase and improve confidence interval reporting in the field, this paper evaluates confidence interval methods for performance metrics under conditions typical of social science text classification: small to moderate sample sizes, infrequent constructs, and texts nested within individuals. Across simulations, default methods such as the Wald interval and the basic percentile bootstrap are the least accurate, with coverage sometimes far below the nominal 95% level. Accuracy is improved with the use of Agresti-Coull, Wilson, Clopper-Pearson, and a novel pseudo-count regularized bootstrap (which is particularly relevant to the calculation of F1). When texts are nested within individuals, we demonstrate that adjustment for both effective N and the appropriate degrees of freedom is necessary for producing accurate analytic intervals. Among bootstrap intervals, the hierarchical bootstrap is more accurate than the cluster bootstrap when individuals produce a moderate number of texts but overly conservative when individuals produce only a few. By providing guidance to the field on appropriate interval estimation, we aim to improve the transparency of machine learning applications, and to encourage greater attention to the validation sample size at the design stage.

文本分类不确定性估计置信区间大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。