用稀疏自编码器揭示大模型与人脑语义的对应关系
Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography
- 用稀疏自编码器分解大模型,提取可解释的语义特征
- 94%的脑响应预测性能由语义特征实现,显著优于基线
- 首次在细粒度上验证语义特征与大脑区域的对应关系
大型语言模型(LLMs)中间层能最好地预测人脑对语言的反应,这是计算神经语言学中最稳健的现象之一,但其机制仍不明确。本文通过将稀疏自编码器(SAEs)与神经编码模型结合,将GPT-2 XL和Llama-3.1-8B分解为每层16K–32K个可解释特征。经人类验证的语义分类体系(κ≥0.74)显示,仅靠语义特征即可恢复94%的峰值编码性能(r=0.285),远超方差匹配基线(p<0.001,d=1.31)。进一步测试了基于三个独立神经科学研究先验的五个语义子类是否应映射到特定脑区,统计检验确认该对应关系(斯皮尔曼ρ=0.72,p<0.001;超几何检验p=0.007),表明SAE发现的特征复现了已知的大脑语义拓扑结构,精度超越以往方法。这些特征还能在控制词汇因素后预测人类阅读时长(ΔlogLik=38.4,p<0.001),初步分析提示大脑可能还编码意外语义内容。结果在英语、中文和法语中均成立。
原文摘要 · Abstract (English)
Intermediate layers of large language models (LLMs) best predict human brain responses to language, one of the most robust findings in computational neurolinguistics, yet why remains mechanistically unexplained. We address this gap by bridging sparse autoencoders (SAEs) from mechanistic interpretability with neural encoding models, decomposing GPT-2 XL and Llama-3.1-8B into 16K-32K interpretable features per layer. A human-validated taxonomy ($κ\geq 0.74$) reveals that semantic features alone recover 94% of peak encoding performance ($r=0.285$), substantially exceeding variance-matched baselines ($p<0.001$, $d=1.31$). Beyond this aggregate dominance, we test a novel cortical topography prediction: five semantic subcategories derived a priori from three independent neuroscience programs should map onto distinct brain regions. A formal convergence test confirms this alignment (Spearman $ρ=0.72$, $p<0.001$; hypergeometric $p=0.007$), demonstrating that SAE-discovered features recapitulate known cortical semantic organization at a granularity inaccessible to prior methods. SAE features further predict human reading times beyond lexical controls ($Δ\mathrm{logLik}=38.4$, $p<0.001$), and an exploratory prediction-error analysis provides preliminary evidence that the brain additionally encodes unexpected semantic content. Results generalize across English, Chinese, and French.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。