通过伪造引用检测大模型幻觉,发现幻觉时隐藏状态呈独特马蹄形。
Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering
- 构建FalseCite数据集,用虚假引用诱导模型产生幻觉。
- GPT-4o-mini在虚假引用下幻觉率显著上升,达37%以上。
- 模型幻觉时隐藏状态呈现统一的马蹄形分布,可作诊断依据。
大型语言模型常产生幻觉,生成错误或无意义信息,在医疗、法律等敏感领域尤为危险。为系统研究此现象,我们提出FalseCite数据集,专门捕捉由误导性或虚构引文引发的幻觉响应。对GPT-4o-mini、Falcon-7B和Mistral 7-B进行测试,发现带有欺骗性引用的虚假陈述导致幻觉行为明显增加,尤其在GPT-4o-mini中表现突出。基于FalseCite的响应结果,我们分析了模型内部状态,可视化并聚类隐藏状态向量。结果显示,无论是否产生幻觉,隐藏状态向量均呈现出独特的马蹄形轨迹。本工作表明FalseCite可作为未来评估与缓解幻觉的重要基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) often hallucinate, generating nonsensical or false information that can be especially harmful in sensitive fields such as medicine or law. To study this phenomenon systematically, we introduce FalseCite, a curated dataset designed to capture and benchmark hallucinated responses induced by misleading or fabricated citations. Running GPT-4o-mini, Falcon-7B, and Mistral 7-B through FalseCite, we observed a noticeable increase in hallucination activity for false claims with deceptive citations, especially in GPT-4o-mini. Using the responses from FalseCite, we can also analyze the internal states of hallucinating models, visualizing and clustering the hidden state vectors. From this analysis, we noticed that the hidden state vectors, regardless of hallucination or non-hallucination, tend to trace out a distinct horn-like shape. Our work underscores FalseCite's potential as a foundation for evaluating and mitigating hallucinations in future LLM research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。