arXiv:2508.09363cs.LG2025-08被引 3

让AI模型理解医学文本更清晰:用专业数据训练稀疏自编码器。

Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders

  • 在医疗文本上训练稀疏自编码器,聚焦特定领域特征。
  • 比通用数据训练的模型多解释20%方差,残差误差更小。
  • 学出的特征贴近临床概念,适合医疗AI可解释性研究。

稀疏自编码器(SAEs)能将大语言模型(LLM)激活分解为潜在特征,揭示其机制结构。传统SAE在广泛数据上训练,受限于固定潜空间预算,仅捕捉高频通用模式,导致重建误差中存在大量线性‘暗物质’,且潜在特征相互混淆或吸收,影响可解释性。本文通过将SAE训练限定在明确领域(医疗文本),重新分配容量以聚焦领域特异性特征,显著提升重建保真度与可解释性。在Gemma-2模型第20层激活上,使用19.5万条临床问答数据训练JumpReLU SAE,结果表明,领域限定型SAE可解释方差比广域型高20%,损失恢复率更高,线性残差误差更小。自动与人工评估均证实,学习到的特征与临床有意义概念(如“味觉感受”或“传染性单核细胞增多症”)对齐,而非频繁但无信息的词元。这些领域特异性SAE捕获了相关线性结构,留下更小、更纯非线性残差。结论表明,领域限定能有效缓解广域SAE的关键缺陷,实现更完整、可解释的潜在分解,提示该领域或需重新审视面向通用目的的‘基础模型’缩放策略。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) decompose large language model (LLM) activations into latent features that reveal mechanistic structure. Conventional SAEs train on broad data distributions, forcing a fixed latent budget to capture only high-frequency, generic patterns. This often results in significant linear ``dark matter'' in reconstruction error and produces latents that fragment or absorb each other, complicating interpretation. We show that restricting SAE training to a well-defined domain (medical text) reallocates capacity to domain-specific features, improving both reconstruction fidelity and interpretability. Training JumpReLU SAEs on layer-20 activations of Gemma-2 models using 195k clinical QA examples, we find that domain-confined SAEs explain up to 20\% more variance, achieve higher loss recovery, and reduce linear residual error compared to broad-domain SAEs. Automated and human evaluations confirm that learned features align with clinically meaningful concepts (e.g., ``taste sensations'' or ``infectious mononucleosis''), rather than frequent but uninformative tokens. These domain-specific SAEs capture relevant linear structure, leaving a smaller, more purely nonlinear residual. We conclude that domain-confinement mitigates key limitations of broad-domain SAEs, enabling more complete and interpretable latent decompositions, and suggesting the field may need to question ``foundation-model'' scaling for general-purpose SAEs.

可解释性稀疏自编码器医疗AI机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。