新知识训练会引发大模型幻觉,且影响会扩散到其他任务。
Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation
- 构建生物人物推理数据集,细粒度分析新知识幻觉
- 陌生知识类型越强,幻觉越严重,与整体新知识比例无关
- 后期引入旧知识可恢复注意力,有效抑制幻觉
已有研究发现,对大语言模型进行新知识微调会引发事实性幻觉,导致在旧知识评估中出现错误输出。然而,这种幻觉的具体表现及其机制仍不明确。本文设计了受控数据集Biography-Reasoning,针对多种知识类型和两类任务(知识问答与知识推理)开展细粒度分析。结果表明,幻觉不仅严重影响涉及新知识的任务,还会传播至其他评估任务。当微调数据集中某一知识类型完全由新知识构成时,模型幻觉倾向显著上升,说明特定知识类型的陌生程度比整体新知识比例更能驱动幻觉。可解释性分析显示,学习新知识会削弱模型对输入问题中关键实体的关注,导致过度依赖上下文而增加幻觉风险。相反,在训练后期引入少量已知知识可恢复对关键实体的注意力,显著缓解幻觉行为。最后,我们发现被破坏的注意力模式可在语义相似的上下文中传播,促进幻觉跨任务扩散。
原文摘要 · Abstract (English)
Prior works have shown that fine-tuning on new knowledge can induce factual hallucinations in large language models (LLMs), leading to incorrect outputs when evaluated on previously known information. However, the specific manifestations of such hallucination and its underlying mechanisms remain insufficiently understood. Our work addresses this gap by designing a controlled dataset \textit{Biography-Reasoning}, and conducting a fine-grained analysis across multiple knowledge types and two task types, including knowledge question answering (QA) and knowledge reasoning tasks. We find that hallucinations not only severely affect tasks involving newly introduced knowledge, but also propagate to other evaluation tasks. Moreover, when fine-tuning on a dataset in which a specific knowledge type consists entirely of new knowledge, LLMs exhibit elevated hallucination tendencies. This suggests that the degree of unfamiliarity within a particular knowledge type, rather than the overall proportion of new knowledge, is a stronger driver of hallucinations. Through interpretability analysis, we show that learning new knowledge weakens the model's attention to key entities in the input question, leading to an over-reliance on surrounding context and a higher risk of hallucination. Conversely, reintroducing a small amount of known knowledge during the later stages of training restores attention to key entities and substantially mitigates hallucination behavior. Finally, we demonstrate that disrupted attention patterns can propagate across lexically similar contexts, facilitating the spread of hallucinations beyond the original task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。