现有医学影像报告评估方法忽略术语丢失问题,导致模型生成空洞内容却得分高。
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
- 引入词汇级评估框架CAD,检测不同人群的临床术语关联偏移
- 发现确定性解码使临床信息大量消失,而随机采样引入新偏差
- 提出加权关联消减(WAE)指标,揭示模型输出的临床失真现象
在放射科应用视觉-语言模型(VLMs)需超越表面文本相似度的验证指标,以保障临床准确性和人口统计公平性。本文揭示当前评估中的关键盲点:某些解码策略虽带来高令牌重叠分数,实则导致模板坍缩——模型仅生成重复、安全的通用语句,遗漏临床术语。若不解决,将引发指标操纵,使表现优异但临床无用的模型被误判为优质。为此,我们主张采用词汇多样性指标检验生成内容的临床特异性。提出临床关联偏移(CAD)框架,量化不同人口群体间词项关联的变化;加权关联消减(WAE)则聚合这些偏移,衡量各群体临床信号的损失程度。实验表明,确定性解码导致严重语义消减,而随机采样虽提升多样性却可能引入新偏差,促使我们重新思考‘最优’报告的定义。
原文摘要 · Abstract (English)
Reliable deployment of Vision-Language Models (VLMs) in radiology requires validation metrics that go beyond surface-level text similarity to ensure clinical fidelity and demographic fairness. This paper investigates a critical blind spot in current model evaluation: the use of decoding strategies that lead to high aggregate token-overlap scores despite succumbing to template collapse, in which models generate only repetitive, safe generic text and omit clinical terminology. Unaddressed, this blind spot can lead to metric gaming, where models that perform well on benchmarks prove clinically uninformative. Instead, we advocate for lexical diversity measures to check model generations for clinical specificity. We introduce Clinical Association Displacement (CAD), a vocabulary-level framework that quantifies shifts in demographic-based word associations in generated reports. Weighted Association Erasure (WAE) aggregates these shifts to measure the clinical signal loss across demographic groups. We show that deterministic decoding produces high levels of semantic erasure, while stochastic sampling generates diverse outputs but risks introducing new bias, motivating a fundamental rethink of how "optimal" reporting is defined.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。