arXiv:2506.21561cs.CLcs.AI2025-06被引 14

发现大模型判真能力仍有偏见,部分模型对谎言识别差

Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs

  • 对比推理与非推理模型的真假判断能力
  • 推理模型真相偏误更低,但依然高于人类基准
  • 多款先进模型存在逢迎倾向,骗术识别能力弱

尽管大语言模型(LLMs)广泛用于事实核查、内容审核和高风险决策,其作为真理裁判者的性能仍不清晰。本研究开展了迄今最大规模的LLM真伪检测评估,并首次分析推理型模型的表现。我们让8个LLM在多个提示下做出4800次真伪判断,对比了推理与非推理模型。结果显示,推理模型的真相偏误率低于非推理模型,但仍高于人类基准。最令人担忧的是,若干先进模型(o4-mini、GPT-4.1、R1)表现出明显的逢迎倾向,其在真实陈述判断上准确率高,但在虚假陈述识别上表现差,说明仅提升模型能力无法解决根本的真伪检测难题。

原文摘要 · Abstract (English)

Despite their widespread use in fact-checking, moderation, and high-stakes decision-making, large language models (LLMs) remain poorly understood as judges of truth. This study presents the largest evaluation to date of LLMs' veracity detection capabilities and the first analysis of these capabilities in reasoning models. We had eight LLMs make 4,800 veracity judgments across several prompts, comparing reasoning and non-reasoning models. We find that rates of truth-bias, or the likelihood to believe a statement is true, regardless of whether it is actually true, are lower in reasoning models than in non-reasoning models, but still higher than human benchmarks. Most concerning, we identify sycophantic tendencies in several advanced models (o4-mini and GPT-4.1 from OpenAI, R1 from DeepSeek), which displayed an asymmetry in detection accuracy, performing well in truth accuracy but poorly in deception accuracy. This suggests that capability advances alone do not resolve fundamental veracity detection challenges in LLMs.

大模型评测真相检测偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。