提出新评估框架STABLEVAL,让AI系统评价更稳定可靠。
STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

- 建模标注者差异与项目模糊性,生成可信评分
- 在真实数据集上,传统投票法误差翻倍,新方法更稳定
- 适合需要可复现评估的AI研究与产品验收
人类评估仍是现代AI系统的主要标准,但标注者分歧、偏见和变异导致标准多数投票聚合下的系统排名脆弱。多数投票忽略标注者可靠性与项目层面的不确定性,常在不同标注者子集中产生不稳定的比较结果。我们提出STABLEVAL,一种关注分歧的评估框架,通过建模潜在项目正确性与标注者特定混淆模式,生成后验期望项目得分和校准的代理级分数。不同于如Dawid-Skene等标签去噪方法,STABLEVAL专为稳定且带不确定性的系统评估而设计,而非硬标签恢复。我们正式将排名稳定性作为首要评估目标,并分析不同聚合方法如何保留或扭曲标注者行为。在受控合成实验与多个真实世界人工标注基准上,多数投票在标注者异质性和对抗性噪声下表现出更高的分数误差与排名不稳定性,而STABLEVAL则产生更稳定且统计稳健的系统排名。结果表明,建模分歧对鲁棒、可复现的AI评估至关重要。
原文摘要 · Abstract (English)
Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majority vote aggregation. Majority vote discards annotator reliability and item-level ambiguity, often yielding unstable comparisons across annotator subsets. We introduce STABLEVAL, a disagreement-aware evaluation framework that models latent item correctness and annotator-specific confusion patterns to produce posterior expected item credit and calibrated agent-level scores. Unlike label-denoising approaches such as Dawid-Skene, STABLEVAL is explicitly designed for stable and uncertainty-aware system evaluation rather than hard label recovery. We formalize ranking stability as a first-class evaluation objective and analyze how aggregation methods preserve or distort underlying annotator behavior. Across controlled synthetic experiments and multiple real-world human-annotated benchmarks, majority vote exhibits increasing score error and ranking instability under annotator heterogeneity and adversarial noise, while STABLEVAL yields more stable and statistically grounded system rankings. These results demonstrate that modeling disagreement is essential for robust and reproducible AI evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。