arXiv:2606.30201cs.CVcs.CL2026-06

提出新基准SHOVIR,检测医学影像报告模型是否依赖视觉捷径

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

论文配图:SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation
图 1 · 摘自论文原文
  • 通过区域遮蔽实验对比模型在完整与局部扰动图像上的表现
  • 发现高分模型仍存在视觉证据缺失时诊断不降反升的现象
  • 适合关注医疗AI可解释性与评估可靠性的研究者参考

当前放射科报告生成任务中,视觉语言模型的评估多依赖于报告层面的词汇重合度或整体临床正确率,无法检验具体诊断结论是否基于图像中实际可见的病灶。这导致模型可能通过利用先验知识或虚假关联获得高分,即存在视觉捷径问题。本文提出SHOVIR基准,扩展了两个带有空间标注的胸片数据集MIMIC-CXR和PadChest-GR,引入每框对应的CheXpert标签,并设计图像级与疾病级遮蔽实验,对比干净图像与局部区域扰动下的模型表现。该方法识别出两类疾病级别故障模式:直接捷径(移除病灶后诊断仍存)与上下文捷径(共现病灶被遮蔽后诊断下降,即使目标区域完好)。对八种前沿视觉语言模型的评测显示,不同架构与数据集间捷径行为差异显著;报告质量最高的模型未必具有强空间对齐能力,说明临床流畅生成可与浅层视觉依赖并存。这一发现揭示现有评估体系的盲点,推动区域感知型评估协议的发展。

原文摘要 · Abstract (English)

Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness. However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image. This allows models to achieve competitive scores by exploiting learned priors or spurious correlations, a failure mode we refer to as vision shortcut. We introduce SHOVIR, a benchmark for evaluating vision shortcut behavior in RRG. SHOVIR extends two spatially annotated chest X-ray datasets, MIMIC-CXR and PadChest-GR, with per-box CheXpert labels, and defines image-level and disease-level occlusion experiments that contrast baseline performance on clean images against localized, region-specific perturbations. Comparing predictions across these conditions isolates two failure modes at the disease-class level: direct shortcuts, where a finding persists after its visual evidence is removed, and contextual shortcuts, where detection degrades once co-occurring pathologies are occluded despite the target region remaining intact. Benchmarking eight state-of-the-art VLMs, we find that shortcut behavior varies substantially across architectures and datasets. Models achieving the highest baseline report quality do not necessarily rank highest in spatial grounding, revealing that clinically fluent generation can coexist with shallow reliance on visual evidence. These findings expose a blind spot in current RRG evaluation and motivate region-aware assessment protocols.

医学影像视觉捷径评估基准可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。