通过白化揭示大模型幻觉的几何本质,区分三类错误类型。
Whitening Reveals Cluster Commitment as the Geometric Separator of Hallucination Types
- 用PCA白化处理嵌入空间,让聚类对齐度成为区分幻觉类型的关键指标。
- 实验显示类型2(高承诺)与类型3(低承诺)在白化后显著分离,类型1居中。
- 发现小样本提示集易引发误判,提示集设计影响微信号分析可靠性。
本文提出一种几何幻觉分类法,将大模型失败分为三类:中心漂移(类型1)、错位收敛(类型2)和覆盖空隙(类型3),依据其在嵌入聚类空间中的特征。先前研究在全维上下文测量中无法区分类型1与类型2。本研究通过对GPT-2-small进行主成分白化与谱分解,结合20次随机种子的多运行稳定性分析及按提示聚合,发现白化后峰值聚类对齐度(max_sim)在霍尔姆校正下可显著区分类型2与类型3,且各类别均值符合预期排序:类型2(最高承诺)> 类型1(中等)> 类型3(最低)。首次在方向稳定但功率不足的条件下观察到类型1/2分离迹象,提示大模型容量限制的可能性。将提示数量从每组15增至30后,原白化熵中的假阳性消失,表明提示集敏感性在近饱和表示空间中显著。谱分解定位该伪影源于主导主成分,且各频段均未出现类型1/2分离,否定谱混叠假设。贡献包括:白化作为揭示聚类承诺为理论正确分离指标的预处理方法;证明类型1/2边界是容量限制而非测量误差;以及提示集在高饱和空间中的脆弱性方法论发现。
原文摘要 · Abstract (English)
A geometric hallucination taxonomy distinguishes three failure types -- center-drift (Type~1), wrong-well convergence (Type~2), and coverage gaps (Type~3) -- by their signatures in embedding cluster space. Prior work found Types~1 and~2 indistinguishable in full-dimensional contextual measurement. We address this through PCA-whitening and eigenspectrum decomposition on GPT-2-small, using multi-run stability analysis (20 seeds) with prompt-level aggregation. Whitening transforms the micro-signal regime into a space where peak cluster alignment (max\_sim) separates Type~2 from Type~3 at Holm-corrected significance, with condition means following the taxonomy's predicted ordering: Type~2 (highest commitment) $>$ Type~1 (intermediate) $>$ Type~3 (lowest). A first directionally stable but underpowered hint of Type~1/2 separation emerges via the same metric, generating a capacity prediction for larger models. Prompt diversification from 15 to 30 prompts per group eliminates a false positive in whitened entropy that appeared robust at the smaller set, demonstrating prompt-set sensitivity in the micro-signal regime. Eigenspectrum decomposition localizes this artifact to the dominant principal components and confirms that Type~1/2 separation does not emerge in any spectral band, rejecting the spectral mixing hypothesis. The contribution is threefold: whitening as preprocessing that reveals cluster commitment as the theoretically correct separating metric, evidence that the Type~1/2 boundary is a capacity limitation rather than a measurement artifact, and a methodological finding about prompt-set fragility in near-saturated representation spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。