arXiv:2604.06756cs.CL2026-04ACL被引 5

研究推理链长度如何影响大模型判断答案真假的准确性

How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality

论文配图:How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality
图 1 · 摘自论文原文
  • 让大模型看推理过程来判断答案对错,但推理流畅度会误导判断
  • 弱模型易被表面流畅的错误推理骗过,强模型也难逃误导
  • 适合关注大模型评估可靠性与推理质量判别的研究者

大语言模型被广泛用作人类评估的可扩展替代品,但这类评估者仍不完善且易受表面偏差影响。一个可能原因是评估者缺乏足够的信息来判断答案正确性。随着具备推理能力的模型兴起,将生成者的推理内容提供给评估者可增加信息量,是提升判断准确性的自然选择。然而其对评估行为的实际影响尚不清楚。本文系统研究了在事实问答和数学推理基准上,访问推理链如何影响基于大模型的判断。发现弱模型极易受推理存在影响,常接受带有流畅推理的错误答案;强模型虽能部分利用推理作为信息证据,但仍会被看似高质量的推理链误导。控制实验进一步表明,推理的流畅性和真实性都是驱动评估决策的关键信号。这些发现强调需要更鲁棒的大模型评估者,能在现代推理模型评估中区分真实推理质量与表面流畅性。

原文摘要 · Abstract (English)

Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One possible reason is that these judges lack sufficient information in assessing answer correctness. With the rise of reasoning-capable models, exposing a generator's reasoning content to the judge provides richer information and is a natural candidate for improving judgment accuracy. However, its actual impact on judge behavior remains understudied. In this paper, we systematically investigate how access to reasoning chains affects LLM-based judgment across factual question answering (QA) and mathematical reasoning benchmarks. We find that weak judges are easily swayed by reasoning presence, frequently accepting incorrect answers accompanied by fluent reasoning, while strong judges can partially leverage reasoning as informative evidence. Nevertheless, even strong judges are misled by seemingly high-quality reasoning chains. Controlled experiments further reveal that both fluency and factuality of reasoning chains are critical signals driving judge decisions. These findings highlight the need for more robust LLM judges that can distinguish genuine reasoning quality from superficial fluency when evaluating modern reasoning models.

大模型评估推理链判断准确性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。