验证几何幻觉分类在GPT-2中的区分能力,发现覆盖间隙最显著。
From Prerequisites to Predictions: Validating a Geometric Hallucination Taxonomy Through Controlled Induction
- 通过控制诱导实验,检验三种几何幻觉分类的有效性。
- 覆盖间隙(Type 3)在静态嵌入中表现稳定(18/20次显著),方向一致但上下文状态下功率不足。
- 类型1和2无法区分,且词级测试因伪重复导致结果虚高。
我们通过在GPT-2中进行受控诱导实验,检验一种几何幻觉分类体系——将失败分为中心偏移(类型1)、错误收敛点(类型2)或覆盖间隙(类型3)——是否具有区分能力。以提示(N=15/组)为推断单元,每项实验运行20次并更换生成种子以量化结果稳定性。在静态嵌入中,类型3的范数分离在18/20次运行中显著(霍尔姆校正后14/20次),中位相关系数r = +0.61;在上下文隐藏状态中,类型3效应方向稳定(19/20次),但因样本量小(N=15)而统计功效不足(仅4/20次显著,中位r = -0.28)。类型1和2在任一空间均未实现有效分离(≤3/20次)。词级测试因伪重复导致显著性虚增4–16倍,该现象在所有20次运行中均复现。结果表明,覆盖间隙是唯一具明显几何特征的幻觉模式,其差异由幅度而非方向驱动,且类型1/2无法区分的现象在124M参数规模下为真实存在。
原文摘要 · Abstract (English)
We test whether a geometric hallucination taxonomy -- classifying failures as center-drift (Type~1), wrong-well convergence (Type~2), or coverage gaps (Type~3) -- can distinguish hallucination types through controlled induction in GPT-2. Using a two-level statistical design with prompts ($N = 15$/group) as the unit of inference, we run each experiment 20 times with different generation seeds to quantify result stability. In static embeddings, Type~3 norm separation is robust (significant in 18/20 runs, Holm-corrected in 14/20, median $r = +0.61$). In contextual hidden states, the Type~3 norm effect direction is stable (19/20 runs) but underpowered at $N = 15$ (significant in 4/20, median $r = -0.28$). Types~1 and~2 do not separate in either space (${\leq}\,3/20$ runs). Token-level tests inflate significance by 4--16$\times$ through pseudoreplication -- a finding replicated across all 20 runs. The results establish coverage-gap hallucinations as the most geometrically distinctive failure mode, carried by magnitude rather than direction, and confirm the Type~1/2 non-separation as genuine at 124M parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。