arXiv:2509.22612cs.CL2025-09中稿 · IJCNLP-AACL 2025

提出新方法量化多语言多任务NLP评估中的不确定性与变异性。

Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation

  • 基于重抽样方法,分离模型与数据带来的性能波动
  • 发现忽略双重来源会严重低估重复实验的变异程度
  • 可评估排行榜中模型排名与差异的可复现性,适合严谨评估者

我们提出一系列基于重抽样的方法,用于量化多语言和/或多任务NLP基准测试中评估指标的不确定性与统计精度。结果显示,性能分数的实验差异源于模型和数据两方面因素,必须同时考虑两者才能避免对假设重复实验的整体变异性产生严重低估。以多语言问答、机器翻译和命名实体识别为例,我们还展示了重抽样方法在量化排行榜中各类量值(如模型排名、模型间差异)的可复现不确定性方面的实用性。

原文摘要 · Abstract (English)

We introduce a set of resampling-based methods for quantifying uncertainty and statistical precision of evaluation metrics in multilingual and/or multitask NLP benchmarks. We show how experimental variation in performance scores arises from both model and data-related sources, and that accounting for both of them is necessary to avoid substantially underestimating the overall variability over hypothetical replications. Using multilingual question answering, machine translation, and named entity recognition as example tasks, we also demonstrate how resampling methods are useful for quantifying the replication uncertainty of various quantities used in leaderboards such as model rankings and pairwise differences between models.

NLP评估不确定性量化多语言可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。