测试视觉语言模型在医学影像中的诊断稳定性,发现高准确率下仍存严重可靠性问题。
Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations

- 通过调换切片顺序和标签位置,检验模型对临床证据不变性的响应
- 序列反转导致48.9%病例预测翻转,标签重排引发67.8%诊断不一致
- 即使移除病灶切片,仍有76.1%病例仍给出确定诊断,暴露过度自信
视觉语言模型(VLM)的标准准确率常掩盖其在敏感领域中的可靠性缺陷。本文利用经病理验证的脑部MRI数据集,系统评估四种VLM家族在保留临床证据的扰动下的诊断鲁棒性。通过重新排列解剖切片顺序及交换目标标签位置,检验模型在临床证据不变时是否保持一致预测。结果表明,模型在呈现顺序稳定性方面存在显著脆弱性,简单序列反转导致高达48.9%的案例出现预测翻转;同时发现文本选择偏差:标签重排引发高达67.8%的案例出现诊断不一致,尽管视觉输入完全相同。负控实验进一步揭示诊断过度承诺现象:在专家标注的病灶切片被移除后,模型仍生成类别化诊断,比例高达76.1%。这些结果说明,高准确率可能夸大临床可靠性,掩盖对序列呈现与文本框架的敏感性,而这类问题无法通过平均准确率捕捉。研究强调,在安全关键的临床应用中部署VLM需引入基于稳定性的评估指标。评估数据与代码将在论文接受后公开。
原文摘要 · Abstract (English)
Standard accuracy metrics for VLMs often mask significant reliability failures in sensitive domains. In this work, we utilize a histopathology-validated brain MRI dataset to systematically assess the diagnostic robustness of four VLM families under evidence-preserving perturbations. By reordering anatomical slices and swapping target label positions, we evaluate whether models maintain consistent predictions when clinical evidence remains invariant. Our results reveal significant vulnerabilities in presentation-order stability, with models exhibiting prediction flips in up to 48.9% of cases under simple sequence reversals. We further identify a textual selection bias, where label reordering triggers inconsistent diagnoses in up to 67.8% of cases despite identical visual inputs. Negative-control tests further reveal diagnostic overcommitment: models generate categorical diagnoses in up to 76.1% of cases after expert-annotated lesion slices are removed. These results demonstrate that high accuracy can overestimate clinical reliability, masking sensitivity to sequential presentation and textual framing that is not captured by aggregate accuracy. Our findings highlight the necessity of stability-based metrics for the deployment of VLMs in safety-critical clinical applications. Our evaluation data and code will be made public upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。