揭示大模型幻觉源于推理偏差而非知识缺失
Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

- 用潜在关键任务模型分析推理路径偏差
- 实验证明幻觉可由统计偏好导致而非缺知识
- 适合研究模型可靠性与因果推理的学者
大型语言模型常产生违反提示约束的幻觉答案。核心问题是这些错误源于知识缺失,还是模型虽有信息却走错了推理路径?我们将其视为推理错位:提示支持的答案与模型依赖的统计显著隐含关联之间存在不一致。通过潜在线性关键任务模型形式化这一观点,发现预训练频率不平衡会导致快捷路径压倒约束敏感路径,引发正向推理损失。该框架预测两种失效模式:实体消歧中的任务-检索偏差,以及动作选择中的关键项选择偏差。我们提出TrapQA,一个包含两部分的受控诊断测试平台:ScientistQA用于测试相似科学家的消歧及补充事实探针,Real-Life Constrained QA则评估日常约束遵循中显性捷径的影响。结果表明,幻觉可源于偏倚的隐式推理,而非仅因知识缺失。
原文摘要 · Abstract (English)
Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures reflect missing knowledge, or whether the model has the relevant information but follows the wrong inference path. We study this phenomenon as inference misalignment: a mismatch between the answer supported by the prompt and the answer favored by statistically salient latent associations. We formalize this view with a latent key-task model, in which pretraining-frequency imbalance can cause a shortcut path to dominate the constraint-sensitive path and induce positive inference loss. The framework predicts two failure modes: task-retrieval bias in entity disambiguation and key-selection bias in action choice. We introduce TrapQA, a controlled diagnostic testbed with two components. ScientistQA tests disambiguation among similar scientists with supplementary factual probes, while Real-Life Constrained QA tests everyday constraint following under salient shortcuts. Our results show that hallucination can arise from biased latent inference rather than absent knowledge alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。