通过嵌入聚类几何结构,首次实现大模型幻觉的可测量分类。
Detecting LLM Hallucinations via Embedding Cluster Geometry: A Three-Type Taxonomy with Measurable Signatures
- 基于嵌入空间聚类几何,提出三类可测幻觉类型
- 11个模型中100%呈现极性耦合与聚类凝聚力
- 揭示模型架构对幻觉敏感性的内在机制
我们基于可观察的词元嵌入聚类结构,提出了大语言模型幻觉的几何分类体系。通过对11种不同架构(包括编码器类BERT、RoBERTa、ELECTRA、DeBERTa、ALBERT、MiniLM、DistilBERT和解码器类GPT-2)的Transformer模型进行静态嵌入空间分析,识别出三类操作上不同的幻觉类型:弱上下文下的中心偏移型(Type 1)、局部一致但语境错误的聚类收敛型(Type 2),以及无聚类结构的覆盖缺口型(Type 3)。引入三种可度量的几何统计量:α(极性耦合)、η(聚类凝聚度)和λ_s(径向信息梯度)。在全部11个模型中,极性结构(α > 0.5)普遍存在(11/11),聚类凝聚度(η > 0)也普遍成立(11/11),径向信息梯度显著(9/11,p < 0.05)。未达显著性的两个模型(ALBERT与MiniLM)其原因可由架构解释:前者因嵌入压缩因子化,后者因蒸馏导致各向同性。该研究确立了特定类型幻觉检测的几何前提,并给出架构依赖性脆弱性预测。
原文摘要 · Abstract (English)
We propose a geometric taxonomy of large language model hallucinations based on observable signatures in token embedding cluster structure. By analyzing the static embedding spaces of 11 transformer models spanning encoder (BERT, RoBERTa, ELECTRA, DeBERTa, ALBERT, MiniLM, DistilBERT) and decoder (GPT-2) architectures, we identify three operationally distinct hallucination types: Type 1 (center-drift) under weak context, Type 2 (wrong-well convergence) to locally coherent but contextually incorrect cluster regions, and Type 3 (coverage gaps) where no cluster structure exists. We introduce three measurable geometric statistics: α (polarity coupling), \b{eta} (cluster cohesion), and λ_s (radial information gradient). Across all 11 models, polarity structure (α > 0.5) is universal (11/11), cluster cohesion (\b{eta} > 0) is universal (11/11), and the radial information gradient is significant (9/11, p < 0.05). We demonstrate that the two models failing λ_s significance -- ALBERT and MiniLM -- do so for architecturally explicable reasons: factorized embedding compression and distillation-induced isotropy, respectively. These findings establish the geometric prerequisites for type-specific hallucination detection and yield testable predictions about architecture-dependent vulnerability profiles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。