医学视觉语言模型常误判无效影像,该研究提出新基准检测其基础判断力。
MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage
- 构建1880任务基准,测试模型对多图集一致性的整体验证能力
- 17个模型在正常图像上仍错误报告异常,大图集下性能显著下降
- 揭示医疗AI安全关键短板:诊断前的图像合理性检查尚未解决
视觉语言模型(VLMs)在医学报告生成和视觉问答中应用日益广泛,但流畅的诊断文本不等于可靠的视觉理解。临床实践中,解读始于预诊断筛查:确认输入是否可读(正确模态与解剖结构、合理视角方向、无明显完整性破坏)。现有基准大多假设此步骤已解决,因而忽略了关键失败模式——模型可在输入不一致或无效时仍生成合理叙述。本文提出MedObvious,一个包含1,880个任务的基准,将输入验证作为小多面板图像集上的集合级一致性能力进行测试:模型需识别是否存在违反预期一致性的图像。MedObvious涵盖五个渐进层级,从基本方位/模态错配到临床驱动的解剖结构/视角验证及分诊式提示,并包含五种评估格式以检验跨接口鲁棒性。对17个不同VLMs的评估显示,预诊断验证仍不可靠:多个模型在正常(负控)输入上虚构异常,规模扩大至更大图像集时性能下降,且多项选择与开放回答设置间准确率差异显著。结果表明,医学VLM的预诊断验证仍未解决,应被视为部署前必须独立处理的安全关键能力。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) are increasingly used for tasks like medical report generation and visual question answering. However, fluent diagnostic text does not guarantee safe visual understanding. In clinical practice, interpretation begins with pre-diagnostic sanity checks: verifying that the input is valid to read (correct modality and anatomy, plausible viewpoint and orientation, and no obvious integrity violations). Existing benchmarks largely assume this step is solved, and therefore miss a critical failure mode: a model can produce plausible narratives even when the input is inconsistent or invalid. We introduce MedObvious, a 1,880-task benchmark that isolates input validation as a set-level consistency capability over small multi-panel image sets: the model must identify whether any panel violates expected coherence. MedObvious spans five progressive tiers, from basic orientation/modality mismatches to clinically motivated anatomy/viewpoint verification and triage-style cues, and includes five evaluation formats to test robustness across interfaces. Evaluating 17 different VLMs, we find that sanity checking remains unreliable: several models hallucinate anomalies on normal (negative-control) inputs, performance degrades when scaling to larger image sets, and measured accuracy varies substantially between multiple-choice and open-ended settings. These results show that pre-diagnostic verification remains unsolved for medical VLMs and should be treated as a distinct, safety-critical capability before deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。