语言模型幻觉源于训练奖励猜测而非承认不确定,统计上是二分类错误。
Why Language Models Hallucinate
- 将幻觉视为二分类错误,由训练与评估机制驱动。
- 不承认不确定会提升测试得分,导致模型持续猜测。
- 需改革现有评测打分方式,而非新增幻觉检测指标。
大型语言模型在不确定时会像考生面对难题一样猜测,产生看似合理却错误的陈述,这种‘幻觉’即使在最先进的系统中也普遍存在,削弱了可信度。我们指出,语言模型产生幻觉的根本原因在于训练和评估过程奖励猜测而非承认不确定性,分析了现代训练流程中幻觉的统计成因。幻觉并非神秘现象,而是二分类中的自然统计误差:当错误陈述与事实无法区分时,预训练语言模型就会在统计压力下产生幻觉。此外,由于大多数评估以得分为导向,模型被优化为擅长应试,不确定时猜测反而能提高表现。这种惩罚不确定响应的‘流行病’只能通过社会技术手段缓解——修改当前主导排行榜但存在偏差的评测评分方式,而非引入额外的幻觉评估。这一变革可能引导领域走向更可信的AI系统。
原文摘要 · Abstract (English)
Like students facing hard exam questions, large language models sometimes guess when uncertain, producing plausible yet incorrect statements instead of admitting uncertainty. Such "hallucinations" persist even in state-of-the-art systems and undermine trust. We argue that language models hallucinate because the training and evaluation procedures reward guessing over acknowledging uncertainty, and we analyze the statistical causes of hallucinations in the modern training pipeline. Hallucinations need not be mysterious -- they originate simply as errors in binary classification. If incorrect statements cannot be distinguished from facts, then hallucinations in pretrained language models will arise through natural statistical pressures. We then argue that hallucinations persist due to the way most evaluations are graded -- language models are optimized to be good test-takers, and guessing when uncertain improves test performance. This "epidemic" of penalizing uncertain responses can only be addressed through a socio-technical mitigation: modifying the scoring of existing benchmarks that are misaligned but dominate leaderboards, rather than introducing additional hallucination evaluations. This change may steer the field toward more trustworthy AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。