大模型幻觉更危险:高自信错误输出难发现难修正
Delusions of Large Language Models
- 识别出高自信错误输出(delusion)这一新型幻觉现象
- 实验显示其在多任务中普遍且难以通过微调或反思消除
- 提出检索增强与多智能体辩论作为缓解策略
大型语言模型常生成看似合理但事实错误的输出,称为幻觉。本文揭示一种更隐蔽的现象——模型妄想(LLM delusion),即高置信度的错误输出,具有异常高的信心但实际不正确,导致检测和缓解更加困难。与普通幻觉不同,妄想表现出低不确定性,严重威胁模型可靠性。通过在多个模型家族和规模上对问答任务的实证分析,我们发现妄想普遍存在且与普通幻觉有本质区别。模型在妄想状态下诚实度更低,且难以通过微调或自我反思纠正。研究将妄想形成归因于训练动态和数据噪声,并探索了检索增强生成与多智能体辩论等缓解方法。本研究系统揭示了妄想的本质、分布及应对策略,为提升模型可靠性提供了关键洞见。
原文摘要 · Abstract (English)
Large Language Models often generate factually incorrect but plausible outputs, known as hallucinations. We identify a more insidious phenomenon, LLM delusion, defined as high belief hallucinations, incorrect outputs with abnormally high confidence, making them harder to detect and mitigate. Unlike ordinary hallucinations, delusions persist with low uncertainty, posing significant challenges to model reliability. Through empirical analysis across different model families and sizes on several Question Answering tasks, we show that delusions are prevalent and distinct from hallucinations. LLMs exhibit lower honesty with delusions, which are harder to override via finetuning or self reflection. We link delusion formation with training dynamics and dataset noise and explore mitigation strategies such as retrieval augmented generation and multi agent debating to mitigate delusions. By systematically investigating the nature, prevalence, and mitigation of LLM delusions, our study provides insights into the underlying causes of this phenomenon and outlines future directions for improving model reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。