arXiv:2606.21203cs.CLcs.AI2026-06

发现大模型也会被语义误导,用三个指标揭示其幻觉机制

When Context Misleads: Surprisal, Energy and Attention Entropy as Metrics of Coherence Illusions in LLMs

论文配图:When Context Misleads: Surprisal, Energy and Attention Entropy as Metrics of Coherence Illusions in LLMs
图 1 · 摘自论文原文
  • 用预测意外度、注意力熵和记忆能量三指标分析模型对连贯性幻觉的响应
  • 当上下文出现匹配词时,模型对不连贯句子的惊讶度显著降低
  • 注意力熵与能量值可识别共享的误导机制,适合研究语言模型认知偏差

心理语言学研究表明,人类读者会陷入连贯性幻觉:即使话语不连贯,只要后续内容与前文存在匹配项(如'again'或'too'),就可能误认为连贯。我们研究了6个荷兰语单语模型和4个多语言模型在类似文本上的表现。结果发现,关键位置的预测意外度(surprisal)与人类可接受度判断及眼动数据高度一致;面对不连贯延续时,模型本应更惊讶,但若前文存在匹配干扰项,惊讶度会下降。注意力熵能识别出在连贯与不连贯条件下行为不同的注意力头,且这些头的移除在不同实验中表现出跨任务转移效应,暗示共享机制存在。我们引入关联记忆文献中的能量概念作为话语连贯性的量化指标。综合来看,荷兰语大模型确实会出现连贯性幻觉,而熵与能量揭示了跨场景运作的内在机制。

原文摘要 · Abstract (English)

Psycholinguistics studies show that human readers fall for coherence illusions: an incoherent discourse can seem coherent simply because a distractor matches what comes next. We investigate whether Dutch language models (6 monolingual and 4 multilingual) show the same behavior on texts that link back to earlier context with words such as 'again' and 'too'. First, we find that surprisal at the critical word tracks human acceptability judgments and eye-tracking data. Models are more surprised by incoherent continuations, but a matching distractor in the prior context reduces this surprisal. Second, attention entropy at the critical position identifies heads that behave differently under coherence vs. incoherence. We find that ablating these heads shows transfer effects across experiments, suggesting a shared mechanism. Third, we introduce energy from the associative-memory literature as a metric to quantify discourse coherence. Taken together, our results show that coherence illusions arise in Dutch LLMs, with entropy and energy exposing mechanisms that operate across settings.

大模型幻觉注意力熵语义连贯性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。