首次用符号语言知识定位大模型幻觉源头,发现否定等符号触发早期崩溃。
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
- 基于符号语言学知识构建新定位框架,聚焦修饰、否定等符号触发机制。
- 早期层(2-4层)注意力方差剧增,否定导致严重不稳定性,幻觉率78.3%-83.7%。
- 揭示幻觉本质是符号语义处理失败,适合研究模型可靠性与对齐的学者。
大语言模型在面对修饰、否定、数字、例外和命名实体等符号触发时仍易产生幻觉,但其根源尚不清晰。现有统计方法如LSC和激活方差分析忽视了符号语言知识的作用,无法系统定位幻觉发生位置。本文提出首个基于符号语言与语义知识的定位框架,分析五种模型在HaluEval和TruthfulQA上的表现。结果表明,符号元素在早期层(2-4层)引发注意力方差剧烈波动,其中否定触发灾难性不稳定;尽管模型规模增大,幻觉率仍维持在78.3%-83.7%之间,深层中符号语义触发的注意力显著下降。研究证明,幻觉本质上是符号语义处理失败,而非泛化生成问题,符号语义知识为理解与定位幻觉机制提供关键路径。
原文摘要 · Abstract (English)
LLMs still struggle with hallucination, especially when confronted with symbolic triggers like modifiers, negation, numbers, exceptions, and named entities. Yet, we lack a clear understanding of where these symbolic hallucinations originate, making it crucial to systematically handle such triggers and localize the emergence of hallucination inside the model. While prior work explored localization using statistical techniques like LSC and activation variance analysis, these methods treat all tokens equally and overlook the role symbolic linguistic knowledge plays in triggering hallucinations. So far, no approach has investigated how symbolic elements specifically drive hallucination failures across model layers, nor has symbolic linguistic knowledge been used as the foundation for a localization framework. We propose the first symbolic localization framework that leverages symbolic linguistic and semantic knowledge to meaningfully trace the development of hallucinations across all model layers. By focusing on how models process symbolic triggers, we analyze five models using HaluEval and TruthfulQA. Our symbolic knowledge approach reveals that attention variance for these linguistic elements explodes to critical instability in early layers (2-4), with negation triggering catastrophic variance levels, demonstrating that symbolic semantic processing breaks down from the very beginning. Through the lens of symbolic linguistic knowledge, despite larger model sizes, hallucination rates remain consistently high (78.3%-83.7% across Gemma variants), with steep attention drops for symbolic semantic triggers throughout deeper layers. Our findings demonstrate that hallucination is fundamentally a symbolic linguistic processing failure, not a general generation problem, revealing that symbolic semantic knowledge provides the key to understanding and localizing hallucination mechanisms in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。