arXiv:2609.06367cs.LGcs.AI2026-09

用多个大模型评估文本生成,给出更可靠的不确定度估计。

Robust Conformal Consensus: Multi-Agent LLM-as-a-Judge Interval Evaluation with Conformal Prediction

  • 融合多个大模型评分构建置信区间,提升评估稳定性。
  • 实验验证区间覆盖率达预期,多模型聚合降低波动。
  • 适合需要可信评估结果的场景,如评测系统设计。

LLM-as-a-Judge 已成为自然语言生成评估的有前景范式。然而,此类评估所伴随的不确定性尚未得到充分探索,限制了其在真实场景中的可靠性。尽管符合性预测(Conformal Prediction)提供了不确定性量化的理论框架,但现有方法通常仅适用于单个 LLM 判官,忽略了不同大模型评估器带来的变异性。本文提出一种面向多智能体 LLM-as-a-Judge 评估的鲁棒不确定性估计框架。该方法基于多个 LLM 的评分构建符合性预测区间,通过综合不同判官的区间输出,获得更稳定可靠的不确定性估计。大量实验证明,所提方法能生成具有覆盖保证的有效预测区间,且跨多个判官的区间聚合显著提升了评估结果的稳定性。

原文摘要 · Abstract (English)

LLM-as-a-Judge has emerged as a promising paradigm for evaluating natural language generation. However, the uncertainty associated with such evaluations remains largely unexplored, which limits their reliability in real-world applications. Although conformal prediction offers a principled framework for uncertainty quantification, existing approaches typically apply it to a single LLM judge, overlooking the variability introduced by using different LLM evaluators. In this work, we propose a robust uncertainty estimation framework for multi-agent LLM-as-a-Judge evaluation. Our approach constructs conformal prediction intervals for LLM-based scores from multiple LLMs. By considering intervals from different LLM judges, we obtain more stable and reliable uncertainty estimates. Extensive experiments demonstrate that our method produces valid prediction intervals with coverage guarantees, and that interval-based aggregation across multiple judges leads to more stable evaluation outcomes.

大模型评估置信区间多模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。