arXiv:2510.26411cs.AI2025-10中稿 · ICIP 2026被引 1

用稀疏自编码器解析医学视觉模型的内部表示,提升可解释性。

MedSAE: Dissecting MedCLIP Representations with Sparse Autoencoders

  • 通过稀疏自编码器分析MedCLIP的隐空间特征。
  • 在CheXpert数据集上实现更高语义单一性和可解释性。
  • 适合关注医疗AI透明性与临床可信性的研究者。

医疗人工智能需要既准确又可解释的模型。本文通过在基于胸部X光片和报告训练的视觉语言模型MedCLIP的隐空间中应用医学稀疏自编码器(MedSAEs),推进了医疗视觉领域的机制可解释性。为量化可解释性,提出结合相关性度量、熵分析及基于MedGemma基础模型的自动神经元命名的评估框架。在CheXpert数据集上的实验表明,MedSAE神经元的单义性与可解释性优于原始MedCLIP特征。研究结果连接了高性能医疗AI与透明性,为构建临床可靠表征提供了可扩展路径。支持本研究发现的源代码可在https://github.com/EIDOSLAB/MedSAE获取。

原文摘要 · Abstract (English)

Artificial intelligence in healthcare requires models that are accurate and interpretable. We advance mechanistic interpretability in medical vision by applying Medical Sparse Autoencoders (MedSAEs) to the latent space of MedCLIP, a vision-language model trained on chest radiographs and reports. To quantify interpretability, we propose an evaluation framework that combines correlation metrics, entropy analyses, and automated neuron naming via the MedGemma foundation model. Experiments on the CheXpert dataset show that MedSAE neurons achieve higher monosemanticity and interpretability than raw MedCLIP features. Our findings bridge high-performing medical AI and transparency, offering a scalable step toward clinically reliable representations. The source code supporting the findings of this study is available at https://github.com/EIDOSLAB/MedSAE.

可解释AI医学图像自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。