arXiv:2605.13801cs.LGcs.AI2026-05

通过多层级标注者建模,提升大模型评估的可复现性

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

  • 构建多层级自举模型,真实模拟标注者行为差异
  • 发现每项任务需至少5次标注才能达统计显著性
  • 适合关注评估可靠性的研究人员和评测平台

随着大语言模型等生成式AI日益普及,确保其安全性、鲁棒性与可信度至关重要。然而,当前AI研究正面临可复现性危机,源于评估不可靠与实验结果无法重复。尽管常使用人工标注评估模型的实用性与安全性,但标注者会引入主观偏见与分歧。现有方法难以克服这一变异,因缺乏足够数据来研究标注池规模增长对实验可复现性的实际影响。标准评估通常每项仅采集3至5次标注,且缺少持久的标注者标识,无法建模个体差异。本文提出一种多层级自举方法,基于拥有大量评分与持久标注者标识的数据集,分析在达成统计显著性时,项目数(N)与每项标注数(K)之间的权衡关系。

原文摘要 · Abstract (English)

As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. However, AI is currently facing a reproducibility crisis driven by unreliable evaluations and unrepeatable experimental results. While human raters are often used to assess models for utility and safety, they introduce divergent biases and subjective opinions into their annotations. Overcoming this variance is exceptionally challenging because very little data exists to study how experimental repeatability actually improves as the annotator pool grows. Standard evaluation practices typically rely on a small number of annotations per item (often 3 to 5) and lack the persistent rater identifiers necessary to model individual variance across items. In this work, we introduce a multi-level bootstrapping approach to model annotator behavior realistically. Leveraging datasets with a large number of ratings and persistent rater identifiers, we analyze the tradeoffs between the number of items ($N$) and the number of responses per item ($K$) required to achieve statistical significance.

可复现性评估方法标注者建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。