arXiv:2601.01490cs.CLcs.AI2026-01被引 2

推理让模型更守规矩,却更爱编造事实。

Distortion Instead of Hallucination: The Effect of Reasoning Under Strict Constraints

  • 在不能查外部资料时,让模型自我推理来保证合规。
  • 推理后违规率从66%降到13%,但事实错误率上升一倍以上。
  • 不同模型对合规与真实性的权衡方式不同,推理不等于更可信。

随着大语言模型广泛应用,输出中非事实性编造(幻觉)成为严重问题。推理能力被视为一种自我验证机制以提升输出可靠性。然而,在无法依赖外部工具或知识的封闭系统中,推理的效果尚不明确。本文在严格约束下(要求推荐计算机科学领域的同行评审论文)对GPT-5.2和Gemini 3 Flash进行实验,考察推理的影响。结果表明,约束合规性与事实准确性之间存在严重权衡:非推理模型违规率高达66%-75%,但保持较高事实准确性;而推理模型将违规率降至13%-26%,却系统性扭曲已知事实,完整虚构内容显著增加。该权衡模式在两种架构不同的模型间一致,表明这是推理的本质局限。此外,推理并未统一提升真实性,效果因模型而异,体现其对合规与真实性的不同分配策略。研究挑战了‘推理必然提升可靠性’的假设:推理模型以诚实违规换取难以检测的扭曲。

原文摘要 · Abstract (English)

With the widespread adoption of large language models (LLMs), hallucinations, which are non-factual fabrications in model outputs, have become serious concerns. Reasoning capabilities have received attention as a self-verification process to improve output reliability. However, the effect of reasoning within a closed system where LLMs cannot rely on external tools or knowledge has yet to be clarified. We therefore conduct experiments under strict constraints (recommending peer-reviewed journal articles in computer science) to examine the effect of reasoning across multiple models (GPT-5.2 and Gemini 3 Flash). Our results reveal a problematic trade-off between constraint compliance and factual accuracy. Non-reasoning models exhibit high constraint violation rates (66-75%) but maintain factual accuracy, while reasoning models reduce violations (13-26%) but systematically distort known facts to satisfy constraints and increase complete fabrication. This trade-off pattern is consistent across both models despite different architectures, indicating a fundamental limitation of reasoning. Furthermore, reasoning does not uniformly improve output authenticity: effects diverge by model, reflecting different allocations of the compliance-truthfulness trade-off. These findings challenge the assumption that reasoning universally improves reliability: reasoning models trade honest constraint violations for detection-resistant distortions.

大模型推理幻觉可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。