arXiv:2511.16842cs.AIcs.CL2025-11NeurIPS被引 14

用统计方法自动发现AI评测中的错误题目,提升评测可靠性。

Fantastic Bugs and Where to Find Them in AI Benchmarks

  • 通过分析模型回答模式,识别偏离预期的异常题目
  • 在9个主流评测中实现最高84%的错误题识别精度
  • 结合大模型初筛,大幅减少人工审查工作量

评测对推动AI发展至关重要,但无效的评测问题常损害其可靠性。手动排查数千道题目既不现实又成瓶颈。本文提出一种系统性评测修订框架,利用回答模式的统计分析,标记可能有问题的题目供专家复核。该方法基于一个常见假设:均值能充分反映模型性能,即测量实验背后存在单一潜在构念,使每道题目的各项统计量具有预期范围。当实际统计值超出预期范围时,该题目更可能是问题题。在九个广泛使用的基准上,该方法引导专家审查,最高可达84%的精确率。此外,我们引入大模型裁判进行首轮筛选,进一步降低人工成本。两者结合,构建了一个高效、可扩展的评测修订体系。

原文摘要 · Abstract (English)

Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasible but also a critical bottleneck for reliable evaluation. In this work, we introduce a framework for systematic benchmark revision that leverages statistical analysis of response patterns to flag potentially invalid questions for further expert review. Our approach builds on a core assumption commonly used in AI evaluations that the mean score sufficiently summarizes model performance. This implies a unidimensional latent construct underlying the measurement experiment, yielding expected ranges for various statistics for each item. When empirically estimated values for these statistics fall outside the expected range for an item, the item is more likely to be problematic. Across nine widely used benchmarks, our method guides expert review to identify problematic questions with up to 84\% precision. In addition, we introduce an LLM-judge first pass to review questions, further reducing human effort. Together, these components provide an efficient and scalable framework for systematic benchmark revision.

AI评测错误检测自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。