arXiv:2602.00906cs.LGcs.AI2026-02被引 2

信息压缩导致幻觉:模型在存不下所有事实时会误信假信息。

Hallucination is a Consequence of Space-Optimality: A Rate-Distortion Theorem for Membership Testing

  • 将事实记忆建模为成员测试,用速率-失真理论分析信息效率。
  • 在有限存储下,最优策略是给非事实高置信度,导致幻觉不可避免。
  • 理论解释了幻觉的本质,也修正了布隆过滤器的内存下限。

大型语言模型常对无推理模式的‘随机事实’以高置信度产生幻觉。本文将此类事实的记忆建模为成员测试问题,统一了布隆过滤器的离散误差与大模型的连续对数损失。在事实稀疏于合理命题空间的设定下,建立了一条速率-失真定理:最优存储效率由事实与非事实得分分布间的最小KL散度决定。该理论表明,在理想训练、完美数据和封闭世界假设下,即使最优训练也无法避免幻觉——受限容量下的信息论最优策略并非回避或遗忘,而是对部分非事实赋予高置信度。我们在合成与真实数据上验证了这一理论,发现幻觉是损毁压缩的自然结果。该定理还重构并精确化了布隆类型过滤器的经典空间下界,确定了此前未解的双向过滤器加性常数。

原文摘要 · Abstract (English)

Large language models often hallucinate with high confidence on "random facts" that lack inferable patterns. We formalize the memorization of such facts as a membership testing problem, unifying the discrete error metrics of Bloom filters with the continuous log-loss of LLMs. By analyzing this problem in the regime where facts are sparse in the universe of plausible claims, we establish a rate-distortion theorem: the optimal memory efficiency is characterized by the minimum KL divergence between score distributions on facts and non-facts. This theoretical framework provides a distinctive explanation for hallucination under an idealized setting: even with optimal training, perfect data, and a simplified ``closed world'' setting, the information-theoretically optimal strategy under limited capacity is not to abstain or forget, but to assign high confidence to some non-facts, resulting in hallucination. We validate this theory empirically on both synthetic and real-world data, showing that hallucinations persist as a natural consequence of lossy compression. The same theorem recovers and sharpens classical space lower bounds for Bloom-type filters, pinning down an additive constant left open for two-sided filters.

幻觉机制信息论大模型压缩成员测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。