提出单一语义特征提取框架,让医学语言模型解释更稳定可信。
A Monosemantic Attribution Framework for Stable Interpretability in Clinical Neuroscience Transformer-Based Language Models
- 通过构建单义嵌入空间,融合归因与机制解释方法。
- 降低不同解释方法间差异,获得稳定输入重要性评分。
- 适合临床神经科学中阿尔茨海默病早期诊断的可信赖分析。
在阿尔茨海默病进展诊断等临床场景中,语言模型的可解释性仍是关键挑战,早期且可信的预测至关重要。现有归因方法因Transformer模型表示的多义性,导致解释结果高度依赖方法且不稳定;而机制解释方法缺乏与输入输出的直接对齐,也未提供明确重要性分数。本文提出统一的可解释性框架,通过在Transformer层级构建单义嵌入空间,并优化以显式减少方法间差异,生成稳定的输入级重要性评分,同时通过目标层的解压缩表示突出显著特征,推动语言模型在认知健康与神经退行性疾病中的安全可信应用。
原文摘要 · Abstract (English)
Interpretability remains a key challenge for deploying language models (LM) in clinical settings such as progression diagnosis of Alzheimer disease, where early and trustworthy predictions are essential. Existing attribution methods exhibit high inter-method variability and unstable explanations due to the polysemantic nature of Transformer-Based LM and LLM representations, while mechanistic interpretability approaches lack direct alignment with model inputs and outputs and do not provide explicit importance scores. We introduce a unified interpretability framework that integrates attributional and mechanistic perspectives through monosemantic feature extraction. By constructing a monosemantic embedding space at the level of an transformer-based LM layer and optimizing the framework to explicitly reduce inter-method variability, our approach produces stable input-level importance scores and highlights salient features via a decompressed representation of the layer of interest, advancing the safe and trustworthy application of LMs in cognitive health and neurodegenerative disease.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。