arXiv:2603.03437cs.CV2026-03被引 7

现有医学视觉问答评估只看准确率,易被模型捷径欺骗。

Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning

  • 用真实、空白、乱序图像对比模型表现,检测视觉依赖性。
  • 文本强化学习模型准确率高但几乎不依赖图像,负视觉依赖分达-0.09。
  • 超六成回答声称有视觉依据,但三成以上无实据,适合研究评估漏洞者。

近期研究表明,仅使用文本的强化学习(文本-仅RLVR)在多模态医学视觉问答(VQA)基准上可达到或超过图文强化学习(图文-RLVR)的准确率,暗示当前评估协议无法有效衡量模型对视觉信息的因果依赖。本文提出一种反事实评估框架,在四个医学VQA基准(PathVQA、PMC-VQA、SLAKE、VQA-RAD)中使用真实、空白和随机图像进行测试。除准确率外,引入视觉依赖得分(VRS)、图像敏感度(IS)及幻觉视觉推理率(HVRR),以检测模型在答案不变时仍生成视觉描述的情况。结果显示,文本-仅RLVR虽提升准确率,但导致视觉依赖下降:在PathVQA上获得负VRS(-0.09),即模型在图像错配时表现更优;图文-RLVR虽提高准确率,但整体图像敏感度降至39.8%。在VQA-RAD上,两种方法均达63%准确率,但文本-仅版本在空白图像下仍保持81%性能,而图文-版本图像敏感度仅为29%。模型在68%-74%的回答中声称基于视觉,但其中38%-43%为无根据陈述(HVRR)。结果表明,仅以准确率为奖励会诱导模型走捷径,真正进展需依赖显式视觉依赖性的评估与训练机制。

原文摘要 · Abstract (English)

Recent work shows that text-only reinforcement learning with verifiable rewards (RLVR) can match or outperform image-text RLVR on multimodal medical VQA benchmarks, suggesting current evaluation protocols may fail to measure causal visual dependence. We introduce a counterfactual evaluation framework using real, blank, and shuffled images across four medical VQA benchmarks: PathVQA, PMC-VQA, SLAKE, and VQA-RAD. Beyond accuracy, we measure Visual Reliance Score (VRS), Image Sensitivity (IS), and introduce Hallucinated Visual Reasoning Rate (HVRR) to detect cases where models generate visual claims despite producing image-invariant answers. Our findings reveal that RLVR improves accuracy while degrading visual grounding: text-only RLVR achieves negative VRS on PathVQA (-0.09), performing better with mismatched images, while image-text RLVR reduces image sensitivity to 39.8% overall despite improving accuracy. On VQA-RAD, both variants achieve 63% accuracy through different mechanisms: text-only RLVR retains 81% performance with blank images, while image-text RLVR shows only 29% image sensitivity. Models generate visual claims in 68-74% of responses, yet 38-43% are ungrounded (HVRR). These findings demonstrate that accuracy-only rewards enable shortcut exploitation, and progress requires grounding-aware evaluation protocols and training objectives that explicitly enforce visual dependence.

视觉定位医疗AI评估漏洞强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。