发现大模型用统一几何结构表示家族谱系图,提升AI可解释性。
Investigating Representation Universality: Case Study on Genealogical Representations
- 通过锥探针定位残差流中的树状子空间,验证其对问答的因果作用。
- 跨5种模型与不同架构验证,图结构表示具有一致性,损失下降15%以上。
- 适合关注模型可解释性与知识表征的科研人员参考。
为提升可解释性与可靠性,我们研究大语言模型(LLMs)是否使用统一的几何结构来编码离散、图结构的知识。为此,我们提出两项互补实验证据,支持图表示的普遍性。首先,在上下文基因谱系问答任务中,训练一个锥探针以分离残差流激活中的树状子空间,并通过激活修补验证其对相关问题回答的因果影响,结果在五种不同模型上均有效。其次,对不同架构和参数量(OPT、Pythia、Mistral、LLaMA,4.1亿至80亿参数)的模型进行模型拼接实验,通过下一词预测损失的相对退化量化表示对齐程度。由于缺乏图结构的真实表示,研究其表征仍具挑战。深入理解模型表示有助于开发更可解释、鲁棒且可控的AI系统。
原文摘要 · Abstract (English)
Motivated by interpretability and reliability, we investigate whether large language models (LLMs) deploy universal geometric structures to encode discrete, graph-structured knowledge. To this end, we present two complementary experimental evidence that might support universality of graph representations. First, on an in-context genealogy Q&A task, we train a cone probe to isolate a tree-like subspace in residual stream activations and use activation patching to verify its causal effect in answering related questions. We validate our findings across five different models. Second, we conduct model stitching experiments across models of diverse architectures and parameter counts (OPT, Pythia, Mistral, and LLaMA, 410 million to 8 billion parameters), quantifying representational alignment via relative degradation in the next-token prediction loss. Generally, we conclude that the lack of ground truth representations of graphs makes it challenging to study how LLMs represent them. Ultimately, improving our understanding of LLM representations could facilitate the development of more interpretable, robust, and controllable AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。