arXiv:2509.09715cs.CLcs.AI2025-09被引 4

发现大模型幻觉主要源于符号性输入,规模增大也难根治。

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA

  • 通过改写问答格式,定位符号属性是幻觉主因
  • 2B模型幻觉率79%,27B降至63.9%,但修饰词仍超84%
  • 适用于关注模型可信度与输入敏感性的研究者

大型语言模型(LLMs)的幻觉问题已被广泛研究,但其内在脆弱性来源尚未明确。本研究识别并刻画了关键属性,揭示模型内部机制中的漏洞。通过在HaluEval和TruthfulQA两个标准数据集上,将原有问答格式转化为多种变体,缩小符号属性作为幻觉成因的范围。结果表明,Gemma-2-2B在任务与数据集上的平均幻觉率达79.0%;模型规模扩大后,幻觉率下降至Gemma-2-9B的73.6%和Gemma-2-27B的63.9%,整体降低15个百分点。尽管如此,由符号属性引发的幻觉仍大量存在:修饰词幻觉率在84.76%至94.98%之间,命名实体在83.87%至93.96%之间,跨越所有模型与数据集。这表明,无论模型规模如何,符号元素仍严重干扰模型处理,暴露其根本缺陷。

原文摘要 · Abstract (English)

Hallucination in Large Language Models (LLMs) is a well studied problem. However, the properties that make LLM intrinsically vulnerable to hallucinations have not been identified and studied. This research identifies and characterizes the key properties, allowing us to pinpoint vulnerabilities within the model's internal mechanisms. To solidify on these properties, we utilized two established datasets, HaluEval and TruthfulQA and convert their existing format of question answering into various other formats to narrow down these properties as the reason for the hallucinations. Our findings reveal that hallucination percentages across symbolic properties are notably high for Gemma-2-2B, averaging 79.0% across tasks and datasets. With increased model scale, hallucination drops to 73.6% for Gemma-2-9B and 63.9% for Gemma-2-27B, reflecting a 15 percentage point reduction overall. Although the hallucination rate decreases as the model size increases, a substantial amount of hallucination caused by symbolic properties still persists. This is especially evident for modifiers (ranging from 84.76% to 94.98%) and named entities (ranging from 83.87% to 93.96%) across all Gemma models and both datasets. These findings indicate that symbolic elements continue to confuse the models, pointing to a fundamental weakness in how these LLMs process such inputs--regardless of their scale.

幻觉分析符号属性Gemma可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。