arXiv:2604.02685cs.LGcs.AI2026-04被引 1

发现大模型内部存在类信念几何结构,可能揭示其认知表征机制。

Finding Belief Geometries with Sparse Autoencoders

  • 用稀疏自编码器与聚类方法搜索模型表示空间中的单纯形结构子空间。
  • 在Gemma-2-9B中找到13个候选单纯形簇,其中5个通过预测有效性验证。
  • 首次实现被动预测与主动干预结果一致,支持真实信念编码假说。

理解内部表征的几何结构是机制可解释性的核心目标。已有研究显示,训练于隐马尔可夫模型生成序列的Transformer会在残差流中编码概率信念状态,呈现单纯形几何结构,顶点对应潜在生成状态。而大规模语言模型在自然文本上训练是否产生类似结构仍是未解问题。本文提出一套发现候选单纯形结构子空间的流程,结合稀疏自编码器(SAEs)、k-子空间聚类及AANet拟合。在已知信念几何的多部分隐马尔可夫模型上验证该流程后,应用于Gemma-2-9B,识别出13个优先簇(K ≥ 3)。关键挑战在于区分真实的信念编码与拼贴伪影:潜在变量可构成单纯形子空间,但混合坐标未必携带预测信号。因此采用重心预测作为主要判别测试。在13个簇中,3个在近顶点样本上表现显著(Wilcoxon p < 10⁻¹⁴),4个在单纯形内部样本上表现显著。共5个独立真实簇至少通过一项检验,而零假设簇均未通过。其中一个簇768_596在全数据集中达到最高因果操控得分,是唯一同时满足被动预测与主动干预一致的案例。这些结果初步表明,Gemma-2-9B中存在真实的类信念几何结构,并指明了确认该解释所需的结构化评估路径。

原文摘要 · Abstract (English)

Understanding the geometric structure of internal representations is a central goal of mechanistic interpretability. Prior work has shown that transformers trained on sequences generated by hidden Markov models encode probabilistic belief states as simplex-shaped geometries in their residual stream, with vertices corresponding to latent generative states. Whether large language models trained on naturalistic text develop analogous geometric representations remains an open question. We introduce a pipeline for discovering candidate simplex-structured subspaces in transformer representations, combining sparse autoencoders (SAEs), $k$-subspace clustering of SAE features, and simplex fitting using AANet. We validate the pipeline on a transformer trained on a multipartite hidden Markov model with known belief-state geometry. Applied to Gemma-2-9B, we identify 13 priority clusters exhibiting candidate simplex geometry ($K \geq 3$). A key challenge is distinguishing genuine belief-state encoding from tiling artifacts: latents can span a simplex-shaped subspace without the mixture coordinates carrying predictive signal beyond any individual feature. We therefore adopt barycentric prediction as our primary discriminating test. Among the 13 priority clusters, 3 exhibit a highly significant advantage on near-vertex samples (Wilcoxon $p < 10^{-14}$) and 4 on simplex-interior samples. Together 5 distinct real clusters pass at least one split, while no null cluster passes either. One cluster, 768_596, additionally achieves the highest causal steering score in the dataset. This is the only case where passive prediction and active intervention converge. We present these findings as preliminary evidence that genuine belief-like geometry exists in Gemma-2-9B's representation space, and identify the structured evaluation that would be required to confirm this interpretation.

可解释性几何结构大模型信念编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。