首个针对胸部X光视觉语言解释的系统性评测,揭示模型在小病灶上定位能力不足。
XBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography
- 基于注意力与相似度生成可视化解释,评估模型对病灶区域的定位准确性。
- 小或弥散病灶的定位准确率显著下降,整体接地能力与识别能力强相关。
- 专用于胸片预训练的模型接地效果更好,适合医疗AI可解释性研究者使用。
视觉语言模型(VLMs)在医学图像理解中展现出出色的零样本性能,但其文本与图像之间的对齐能力——即文本概念是否准确对应视觉证据——仍缺乏系统评估。在医疗领域,可靠的对齐能力对可解释性和临床应用至关重要。本文首次构建了针对七种基于CLIP的VLM变体在胸部X光图像中跨模态可解释性的系统性基准。通过交叉注意力和基于相似度的定位图生成视觉解释,并量化评估其与放射科医生标注区域的一致性。分析表明:(1)所有模型在大而明确的病灶上表现合理,但在小或弥散病灶上性能明显下降;(2)在胸部X光数据上预训练的模型比通用数据预训练模型具有更优的对齐表现;(3)模型的整体识别能力与接地能力高度相关。结果表明,当前VLM虽识别能力强,但在临床可靠接地方面仍存短板,强调部署前需建立针对性可解释性评估基准。XBench代码已开源。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have recently shown remarkable zero-shot performance in medical image understanding, yet their grounding ability, the extent to which textual concepts align with visual evidence, remains underexplored. In the medical domain, however, reliable grounding is essential for interpretability and clinical adoption. In this work, we present the first systematic benchmark for evaluating cross-modal interpretability in chest X-rays across seven CLIP-style VLM variants. We generate visual explanations using cross-attention and similarity-based localization maps, and quantitatively assess their alignment with radiologist-annotated regions across multiple pathologies. Our analysis reveals that: (1) while all VLM variants demonstrate reasonable localization for large and well-defined pathologies, their performance substantially degrades for small or diffuse lesions; (2) models that are pretrained on chest X-ray-specific datasets exhibit improved alignment compared to those trained on general-domain data. (3) The overall recognition ability and grounding ability of the model are strongly correlated. These findings underscore that current VLMs, despite their strong recognition ability, still fall short in clinically reliable grounding, highlighting the need for targeted interpretability benchmarks before deployment in medical practice. XBench code is available at https://github.com/Roypic/Benchmarkingattention
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。