MATEX提升医学视觉语言模型解释性,让AI诊断更精准可信。
MATEX: Multi-scale Attention and Text-guided Explainability of Medical Vision-Language Models
- 融合多尺度注意力与文本引导的空间先验,生成精准定位的解释图。
- 在MS-CXR数据集上,空间精度和专家标注对齐度均优于当前最佳方法。
- 适合医疗AI可解释性研究者及放射科医生,增强临床信任感。
我们提出MATEX(多尺度注意力与文本引导可解释性)框架,通过引入解剖学相关的空间推理,推进医学视觉语言模型的可解释性。MATEX结合多层注意力传播、文本引导的空间先验以及层级一致性分析,生成精确、稳定且具有临床意义的梯度归因图。针对以往方法存在的空间不精确、缺乏解剖学依据及注意力粒度不足等问题,MATEX实现了更忠实的模型解释。在MS-CXR数据集上的评估显示,其空间精度和与专家标注发现的对齐度均优于当前最先进的M2IB方法。结果表明,MATEX在提升放射科AI应用的信任度与透明度方面具有重要潜力。
原文摘要 · Abstract (English)
We introduce MATEX (Multi-scale Attention and Text-guided Explainability), a novel framework that advances interpretability in medical vision-language models by incorporating anatomically informed spatial reasoning. MATEX synergistically combines multi-layer attention rollout, text-guided spatial priors, and layer consistency analysis to produce precise, stable, and clinically meaningful gradient attribution maps. By addressing key limitations of prior methods, such as spatial imprecision, lack of anatomical grounding, and limited attention granularity, MATEX enables more faithful and interpretable model explanations. Evaluated on the MS-CXR dataset, MATEX outperforms the state-of-the-art M2IB approach in both spatial precision and alignment with expert-annotated findings. These results highlight MATEX's potential to enhance trust and transparency in radiological AI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。