arXiv:2507.12950cs.LG2025-07被引 2

用稀疏自编码器解析医学影像大模型的内部认知,发现可解释的临床特征。

Insights into a radiology-specialised multimodal large language model with sparse autoencoders

  • 通过马特里什卡稀疏自编码器分析模型内部表征
  • 识别出导管、心影增大等20+个临床相关概念
  • 适合医疗AI可解释性研究者参考

可解释性能提升AI模型的安全性、透明度与可信度,尤其在医疗决策这种高风险场景中至关重要。机制可解释性,尤其是通过稀疏自编码器(SAEs)方法,为揭示基于Transformer的大模型中的可人类理解特征提供了前景。本研究将马特里什卡-稀疏自编码器(Matryoshka-SAE)应用于专用于放射学的多模态大语言模型MAIRA-2,开展大规模自动化可解释性分析。我们识别出一系列临床相关概念,包括医疗设备(如导管、引流管置入、起搏器存在)、病理性特征(如胸腔积液、心影增大)、纵向变化及文本特征。通过定向控制实验,我们进一步考察这些特征对模型行为的影响,结果显示生成过程具备一定程度的方向性调控能力,但效果参差不齐。研究揭示了实际与方法层面的挑战,但仍为MAIRA-2所学习的内部概念提供了初步洞察,标志着迈向更深层次的机制理解与可解释性的关键一步,并为提升放射科适配型多模态大模型的透明度铺平道路。我们已公开训练好的SAEs及解释结果:https://huggingface.co/microsoft/maira-2-sae。

原文摘要 · Abstract (English)

Interpretability can improve the safety, transparency and trust of AI models, which is especially important in healthcare applications where decisions often carry significant consequences. Mechanistic interpretability, particularly through the use of sparse autoencoders (SAEs), offers a promising approach for uncovering human-interpretable features within large transformer-based models. In this study, we apply Matryoshka-SAE to the radiology-specialised multimodal large language model, MAIRA-2, to interpret its internal representations. Using large-scale automated interpretability of the SAE features, we identify a range of clinically relevant concepts - including medical devices (e.g., line and tube placements, pacemaker presence), pathologies such as pleural effusion and cardiomegaly, longitudinal changes and textual features. We further examine the influence of these features on model behaviour through steering, demonstrating directional control over generations with mixed success. Our results reveal practical and methodological challenges, yet they offer initial insights into the internal concepts learned by MAIRA-2 - marking a step toward deeper mechanistic understanding and interpretability of a radiology-adapted multimodal large language model, and paving the way for improved model transparency. We release the trained SAEs and interpretations: https://huggingface.co/microsoft/maira-2-sae.

可解释性医学AI自编码器多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。