arXiv:2601.20913cs.LGcs.AI2026-01中稿 · ICLR被引 22

用少量人工标注数据校准AI评判者,让不靠谱的评估也可靠。

Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges

  • 通过小样本人工数据估算裁判者真假阳性率,修正统计阈值。
  • 在多个数据集上验证:即使裁判有噪声,仍能保证错误率控制。
  • 适合关注AI安全评估可靠性的研究者和工程团队。

大型语言模型(LLM)的安全性认证——确保失败率低于安全阈值——至关重要但极具挑战。尽管使用“LLM作为裁判”可提升可扩展性,但裁判的不完美、噪声和偏见会破坏统计保证。本文提出“噪声但有效”的假设检验框架,利用少量人工标注的校准集估计裁判的真阳性率(TPR)与假阳性率(FPR),并据此推导出方差修正后的临界阈值,应用于大规模裁判标注数据集。关键在于,该框架在有限样本下理论上保证了第一类错误控制(有效性),区别于预测驱动推断(PPI)。我们证明:在特定条件下,使用噪声裁判的检验比直接评估具有更高统计功效;在Jigsaw评论、仇恨言论及SafeRLHF数据集上的实验验证了理论;进一步揭示了实际方法与理想“奥拉克尔”(完美已知参数)之间的显著性能差距,量化了参数估计的成本。本工作首次系统处理了裁判不完美场景,提供可解释的裁判可靠性诊断,并阐明评估效能如何依赖裁判质量、数据规模与认证水平。

原文摘要 · Abstract (English)

Reliable certification of Large Language Models (LLMs)-verifying that failure rates are below a safety threshold-is critical yet challenging. While "LLM-as-a-Judge" offers scalability, judge imperfections, noise, and bias can invalidate statistical guarantees. We introduce a "Noisy but Valid" hypothesis testing framework to address this. By leveraging a small human-labelled calibration set to estimate the judge's True Positive and False Positive Rates (TPR/FPR), we derive a variance-corrected critical threshold applied to a large judge-labelled dataset. Crucially, our framework theoretically guarantees finite-sample Type-I error control (validity) despite calibration uncertainty. This distinguishes our work from Prediction-Powered Inference (PPI), positioning our method as a diagnostic tool that explicitly models judge behavior rather than a black-box estimator. Our contributions include: (1) Theoretical Guarantees: We derive the exact conditions under which noisy testing yields higher statistical power than direct evaluation; (2) Empirical Validation: Experiments on Jigsaw Comment, Hate Speech and SafeRLHF confirm our theory; (3) The Oracle Gap: We reveal a significant performance gap between practical methods and the theoretical "Oracle" (perfectly known judge parameters), quantifying the cost of estimation. Specifically, we provide the first systematic treatment of the imperfect-judge setting, yielding interpretable diagnostics of judge reliability and clarifying how evaluation power depends on judge quality, dataset size, and certification levels. Together, these results sharpen understanding of statistical evaluation with LLM judges, and highlight trade-offs among competing inferential tools.

大模型评估统计推断安全认证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。