arXiv:2604.03257cs.CLcs.AI2026-04

用约束最大似然法,结合少量人工标注和大模型判断,精准估算大模型失效率。

Robust LLM Performance Certification via Constrained Maximum Likelihood Estimation

论文配图:Robust LLM Performance Certification via Constrained Maximum Likelihood Estimation
图 1 · 摘自论文原文
  • 融合人工标注、大模型判断与领域约束,构建可解释的失效率估计框架。
  • 在不同判断准确率和数据规模下,误差更小、方差更低,优于现有方法。
  • 适合需要高可信度模型安全评估的研究者与工程团队使用。

准确估算大语言模型(LLM)的失效率是其安全部署的前提。然而,当前实践中常面临高昂的人工黄金标准与潜在严重偏倚的自动标注方案(如“大模型作为评判者”)之间的权衡。本文提出一种基于约束最大似然估计(Constrained MLE)的新方法,集成三类信号:(i) 小规模高质量人工标注校准集,(ii) 大规模大模型评判者标注,以及 (iii) 基于已知评判者性能统计界限的领域特定约束信息。通过全面实证研究验证,该方法在多种实验设置下——涵盖不同评判者准确率、校准集规模及模型失效率——均显著优于当前主流基线(如Prediction-Powered Inference, PPI),提供更准确、方差更小的估计结果。通过超越对自动化评判者的“黑箱”使用,本方法为大模型失效率认证提供了原则性、可解释且可扩展的路径。

原文摘要 · Abstract (English)

The ability to rigorously estimate the failure rates of large language models (LLMs) is a prerequisite for their safe deployment. Currently, however, practitioners often face a tradeoff between expensive human gold standards and potentially severely-biased automatic annotation schemes such as "LLM-as-a-Judge" labeling. In this paper, we propose a new, practical, and efficient approach to LLM failure rate estimation based on constrained maximum-likelihood estimation (MLE). Our method integrates three distinct signal sources: (i) a small, high-quality human-labeled calibration set, (ii) a large corpus of LLM-judge annotations, and, most importantly, (iii) additional side information via domain-specific constraints derived from known bounds on judge performance statistics. We validate our approach through a comprehensive empirical study, benchmarking it against state-of-the-art baselines like Prediction-Powered Inference (PPI). Across diverse experimental regimes -- spanning varying judge accuracies, calibration set sizes, and LLM failure rates -- our constrained MLE consistently delivers more accurate and lower-variance estimates than existing methods. By moving beyond the "black-box" use of automated judges to a flexible framework, we provide a principled, interpretable, and scalable pathway towards LLM failure-rate certification.

大模型评估可靠性统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。