arXiv:2604.04418cs.HCcs.AI2026-04

提出可验证错误的新标准,评估大模型解释是否真能帮人判断答案对错。

Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality

  • 定义'错误可验证性',用平衡指标衡量解释能否帮助判断答案正确性
  • 发现模型规模和训练方法无法提升可验证性,需专门干预
  • 提出两种新方法,结合领域信息显著提升解释可信度,适合高风险场景

随着大语言模型在高风险场景中的应用,用户需判断单个回答的正确性,常依赖模型生成的推理链或解释。然而,目前尚无标准衡量这些解释是否真正帮助用户区分正确与错误答案。本文将此概念形式化为错误可验证性,并提出 $v_{ ext{bal}}$ 这一平衡指标,经人类评估者验证,该指标与人工判断高度一致。研究发现,常规方法如后训练、模型缩放,以及部分针对性优化均无法提升可验证性。为此,我们提出两种有效方法:用于数学推理的反射重述(RR)和用于事实问答的预言重述(OR),二者通过引入领域适配的外部信息,显著提升可验证性。结果表明,错误可验证性是响应质量的一个独立维度,不随准确率提升而自然产生,需采用特定领域感知的方法加以改善。

原文摘要 · Abstract (English)

As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or explanations. Yet, no standard measure exists for whether these justifications help users distinguish correct answers from incorrect ones. We formalize this idea as error verifiability and propose $v_{\text{bal}}$, a balanced metric that measures whether justifications enable raters to accurately assess answer correctness, validated against human raters who show high agreement. We find that neither common approaches, such as post-training and model scaling, nor more targeted interventions recommended improve verifiability. We introduce two methods that succeed at improving verifiability: reflect-and-rephrase (RR) for mathematical reasoning and oracle-rephrase (OR) for factual QA, both of which improve verifiability by incorporating domain-appropriate external information. Together, our results establish error verifiability as a distinct dimension of response quality that does not emerge from accuracy improvements alone and requires dedicated, domain-aware methods to address.

大模型评估可解释性推理验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。