arXiv:2512.08121cs.LGcs.AI2025-12Conference of the …被引 4

用平衡准确率选大模型评判者,避免误判偏差。

Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic

  • 用Youden's J统计量证明平衡准确率更适合作为评判者选择标准
  • 在类别不平衡下,传统指标易误导,平衡准确率更稳定可靠
  • 适合需要公平评估模型行为的场景,如政策合规性检测

大语言模型的严谨评估依赖于对理想或不当行为出现率的比较,如任务通过率或违规事件。这些比率由分类器(如LLM作评判者或人工标注)生成,因此分类器的选择直接影响评估可信度。常用的准确率、精确率和F1值对类别不平衡敏感,且受正类设定影响,可能偏好扭曲真实比例的评判者。本文证明Youden's J统计量在理论上与选择最佳评判者一致,而平衡准确率是其线性等价形式。通过理论分析与实证案例及模拟实验,表明采用平衡准确率可实现更优、更稳健的评判者筛选。

原文摘要 · Abstract (English)

Rigorous evaluation of large language models (LLMs) relies on comparing models by the prevalence of desirable or undesirable behaviors, such as task pass rates or policy violations. These prevalence estimates are produced by a classifier, either an LLM-as-a-judge or human annotators, making the choice of classifier central to trustworthy evaluation. Common metrics used for this choice, such as Accuracy, Precision, and F1, are sensitive to class imbalance and to arbitrary choices of positive class, and can favor judges that distort prevalence estimates. We show that Youden's $J$ statistic is theoretically aligned with choosing the best judge to compare models, and that Balanced Accuracy is an equivalent linear transformation of $J$. Through both analytical arguments and empirical examples and simulations, we demonstrate how selecting judges using Balanced Accuracy leads to better, more robust classifier selection.

大模型评估评价指标分类器选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。