arXiv:2607.08065cs.AI2026-07被引 3

大模型一致时未必正确,一致性仅在特定场景下可作可信度参考

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

  • 通过53个运行者生成26.5万份样本,检验模型间及自一致性是否代表正确性
  • 一致性相关系数仅0.20-0.59,前沿模型虽高一致却更易犯错
  • 适用于中等性能模型和算力分配,不适用于顶尖模型的自信判断

LLM作为评判者(LLM-as-judge)正成为企业级AI评估的主流,常以集成或专家混合方式部署。这类系统依赖一个核心假设:判别结果的一致性(模型自身或不同模型间)意味着正确性。我们揭示该假设不可靠:一致性可能源于共同偏差、记忆化策略或选项位置偏好,而非真实正确。在大规模跨运行者研究中,53个运行者对GPQA Diamond和AIME中的重叠案例生成K=50份样本,共26.5万条输出。以多数正确为部署标签,采用分层运行者聚类自助法分析发现,一致性是正向但弱预测信号(皮尔逊相关系数rho 0.20–0.59,所有情况均为正),其有效性取决于使用场景:最适用于未饱和的中等性能模型与算力分配;最差则出现在最一致的前沿模型——其一致性≥0.8的案例占77%,其中48%错误。对三类Claude模型的交叉检查显示相同现象,高自信错误在多家提供商间重复出现,超过边际保持的零模型基准。因此,自一致性仅为条件性正确代理,不可作为独立置信度指标。数据已公开发布。

原文摘要 · Abstract (English)

LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or "mixture-of-experts" (Shazeer et al., 2017) panels of judges. These systems share a key assumption: that consistency -- agreement among judges, or among a model's own samples -- indicates correctness. We show this assumption is unreliable. Agreement is not accuracy: a model can agree with itself, and different models can agree with each other, out of shared bias, a memorized heuristic, or an option-position prior rather than truth. We ask when agreement is nonetheless a usable proxy, in a large-scale cross-runner study: 53 runners drew K=50 samples for assigned overlapping cases across comparisons of model tier, prompting, and scale on GPQA Diamond and AIME -- 265,000 samples. Using majority-correctness as the deployment label and a hierarchical runner-clustered bootstrap, agreement is a positive but weak predictor (rho 0.20-0.59, all positive under item-clustered resampling) whose usefulness is regime-dependent: best for unsaturated mid-tier models and for allocating compute, and worst -- over-confident yet no more accurate -- for the most consistent frontier model (agreement >=0.8 on 77% of GPQA case-result entries, 48% of those wrong). An exploratory cross-family check on three Claude tiers shows the same frontier over-confidence, with confident errors recurring across providers above a marginal-preserving null. Self-consistency is thus a conditional proxy for correctness, not a standalone confidence score. We publicly release the de-identified per-run rows and answer distributions.

大模型评估一致性置信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。