揭示医学大模型如何存储与处理疾病、药物等知识,指导模型优化。
Medical Interpretability and Knowledge Maps of Large Language Models
- 用四种方法分析模型中间层激活,绘制医学知识分布图。
- 发现医学知识主要集中在Llama3.3-70B前半层,药物按专科聚类更明显。
- 揭示年龄编码非线性、疾病进展呈循环模式,适合医疗领域研究者参考。
我们系统研究了大语言模型(LLMs)在医学领域的可解释性。通过四种可解释性技术:(1) 中间激活的UMAP投影,(2) 基于梯度的显著性分析,(3) 层移除/破坏实验,(4) 激活修补,我们揭示了五种LLMs中患者年龄、症状、疾病和药物知识的存储位置。特别地,在Llama3.3-70B中,大部分医学知识集中在前半部分的模型层。此外,我们观察到几个有趣现象:(i) 年龄在中间层以非线性甚至不连续方式编码;(ii) 疾病进展在某些层呈现非单调且循环的表示;(iii) Llama3.3-70B中,药物按医学专科聚类优于作用机制;(iv) Gemma3-27B和MedGemma-27B在中间层激活坍缩,但在最终层恢复。这些结果可为医学任务中的微调、去偏或知识删减提供层级指导。
原文摘要 · Abstract (English)
We present a systematic study of medical-domain interpretability in Large Language Models (LLMs). We study how the LLMs both represent and process medical knowledge through four different interpretability techniques: (1) UMAP projections of intermediate activations, (2) gradient-based saliency with respect to the model weights, (3) layer lesioning/removal and (4) activation patching. We present knowledge maps of five LLMs which show, at a coarse-resolution, where knowledge about patient's ages, medical symptoms, diseases and drugs is stored in the models. In particular for Llama3.3-70B, we find that most medical knowledge is processed in the first half of the model's layers. In addition, we find several interesting phenomena: (i) age is often encoded in a non-linear and sometimes discontinuous manner at intermediate layers in the models, (ii) the disease progression representation is non-monotonic and circular at certain layers of the model, (iii) in Llama3.3-70B, drugs cluster better by medical specialty rather than mechanism of action, especially for Llama3.3-70B and (iv) Gemma3-27B and MedGemma-27B have activations that collapse at intermediate layers but recover by the final layers. These results can guide future research on fine-tuning, un-learning or de-biasing LLMs for medical tasks by suggesting at which layers in the model these techniques should be applied.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。