arXiv:2605.24492cs.CV2026-05

构建医疗视觉语言模型的对抗性评测基准,检验其是否真正基于视觉证据推理。

Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs

论文配图:Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs
图 1 · 摘自论文原文
  • 设计分阶段问答任务,评估模型在临床流程中每步的视觉证据依赖性。
  • 14个模型在42,432张图像上表现逐阶段下降,对抗样本下准确率显著降低。
  • 适合关注医疗AI可解释性与鲁棒性的研究者使用。

视觉语言模型在通用医学视觉问答中表现优异,但因可解释性不足,其预测是否基于真实临床推理仍不明确。本文提出Med-R2 Bench,一个与临床工作流程对齐的分层评测基准,用于评估模型在视觉证据上的对抗鲁棒性。通过设计分步问答任务,评估模型在四个临床阶段中推理链是否严格依赖视觉证据,并引入对抗扰动测试其对误导性线索的敏感性。Med-R2包含42,432张图像、31个任务类别和110,406个问答对。在14个VLM上的评估显示,模型性能随临床阶段推进呈递减趋势。对抗实验表明,模型严重依赖正确提示来猜测答案,即使提供明确视觉线索,也难以准确匹配文本描述。最后,我们通过分步微调验证了该层级数据能显著提升推理鲁棒性,揭示其推动基于证据的医疗AI发展的潜力。

原文摘要 · Abstract (English)

Vision-language models have demonstrated impressive capabilities in general medical visual question answering, yet due to limited interpretability, it remains unclear whether their predictions reflect evidence-grounded clinical reasoning or reliance on spurious priors. We introduce Med-R2 Bench, a hierarchical benchmark aligned with the clinical workflow to evaluate adversarial robustness with visual grounding. We design stepwise QA tasks to assess whether reasoning chains are strictly grounded in visual evidence across the four clinical stages, and employ adversarial perturbations to test robustness against misleading cues. Med-R2 comprises 42,432 images, 31 task categories, and 110,406 QA pairs. Evaluation across 14 VLMs reveals a sequential performance degradation along the four-stage clinical workflow. Adversarial experiments show that models rely heavily on correct prompts to guess answers. Even when provided with explicit visual cues, the models struggle to accurately align textual descriptions. Finally, we demonstrate stepwise fine-tuning using our hierarchical data significantly improves reasoning robustness, highlighting its potential to drive future improvements in evidence-based medical AI.

医疗AI视觉语言模型对抗评测可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。