arXiv:2605.07919cs.CV2026-05

测试医学视觉语言模型在图像失真时的可信度,发现多数模型仍会自信回答错误。

MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence

论文配图:MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence
图 1 · 摘自论文原文
  • 构建300例临床真实场景下的对抗性测试集,由四名放射科医生全程标注。
  • 16个模型平均复合得分仅69.2,人类专家达83.3,差距显著。
  • 适用于评估医疗AI在误诊风险场景下的可靠性,适合临床AI开发者使用。

医学视觉-语言模型(VLMs)通常在完整图像-问题对上进行评估,但可信临床应用需更强能力:模型应能识别答案依据失效的情况。本文通过在扰动证据下发生的沉默失败现象来研究此问题,即医学问题需依赖视觉证据,但图像被篡改、文字误导或关键区域受损,模型仍给出流畅非拒绝回答。我们提出MedVIGIL,一个包含300个案例的评估套件,源自四个公开医学VQA数据源,所有标注由四位注册放射科医师端到端监督完成:每条黄金答案、拒答选项、候选答案集、改写句、错误前提陷阱、ROI框及临床风险等级均由临床医生生成。两名主治放射科医师并行标注,资深医师整合发布清单,独立第四位放射科医师对所有探针作答,提供人类基准。数据集包含2556道多选题、240组反事实三元组、医师裁定的风险等级与可回答性标记、ROI框及配对开放式变体。报告七项条件正确性审计指标,汇总为MedVIGIL综合得分(MCS),审计16个视觉模型及两个纯文本基线。独立放射科医师在沉默失败率5.8%下获得MCS 83.3,领先最强模型Claude Opus 4.7(69.2)达14.1分。基准与评估工具已公开。

原文摘要 · Abstract (English)

Medical vision--language models (VLMs) are usually evaluated on intact image--question pairs, but trustworthy clinical use requires a stronger property: a model must recognise when the evidential basis for an answer has failed. We study this through silent failures under perturbed evidence, where a vision-required medical question is paired with a false premise, wording perturbation, knowledge-only rewrite, or ROI-corrupted image, yet the model returns a fluent non-refusal answer. We introduce medvigil, a 300-case evaluation suite drawn from four public medical VQA sources, supervised end to end by four board-certified radiologists: every gold answer, refusal option, candidate-answer set, paraphrase, false-premise trap, ROI box, and clinical risk tier is clinician-authored. Two attending radiologists annotate every case in parallel, a senior radiologist consolidates the released manifest, and a separate fourth radiologist independent of construction answers every probe to provide the human reference baseline. The release contains 2556 MCQ probes, 240 counterfactual triplets, physician-adjudicated risk-tier and answerability flags, ROI boxes, and a paired open-ended variant. We report seven correctness-conditioned audit metrics that summarise into the medvigil Composite Score (MCS), and audit 16 vision-capable models plus two text-only baselines. The independent radiologist scores MCS 83.3 at silent-failure rate 5.8%, leaving a 14.1-point composite headroom above the strongest audited model (Claude Opus 4.7 at 69.2). The benchmark and evaluation harness are publicly released.

医学AI可信度评估视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。