arXiv:2606.30653cs.CYcs.AI2026-06

模型自评时一致性越高,越容易犯错,存在可靠性悖论。

The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes

论文配图:The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes
图 1 · 摘自论文原文
  • 用自一致性指标直接检验模型生成与自评是否一致
  • 491个概念中模型一致性差异大,高一致性反促错误率上升
  • 临床场景下高一致性模型更易出错,适合评估安全风险的开发者关注

大型语言模型越来越多地用于自主代理流程中,依赖模型在无外部验证的情况下自我评估输出。这些流程的可靠性基于一个隐含假设:模型在生成输出和后续评估时对相关概念的运用方式保持一致。我们提出新的衡量指标——生成-评估自一致性,用于直接测试该假设,并在10个前沿模型上针对491个概念进行了评估。结果发现:第一,自一致性存在显著差异;第二,在由医生验证的临床错误数据集(Proniakin et al., 2025)中,自一致性更高的模型反而表现出更大的错误脆弱性。这揭示了大模型中的‘一致性困境’:自一致性虽有助于操作稳定,但高度一致的模型反而更易犯错。因此,仅靠一致性不能保证部署安全性。

原文摘要 · Abstract (English)

Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification. The reliability of these pipelines depends on an implicit assumption: that the model applies relevant concepts the same way when it generates an output and later evaluates that output. We propose a new measure, generator-evaluator self-consistency, to test this assumption directly and apply it to 10 frontier models across 491 concepts. We find, first, that there is substantial variation in self-consistency. Second, we find that in a clinical setting with physician-validated mistakes (Proniakin et al., 2025), across models, those with higher self-consistency are linked to greater vulnerability to mistakes. Thus, even when models consistently apply concepts they may not be safe to deploy. This is evidence of a consistency dilemma in LLMs: self-consistency is operationally useful, but models that are more consistent are also more prone to mistakes.

大模型评估一致性安全风险自检机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。