arXiv:2605.10850cs.CV2026-05被引 2

医学视觉问答中自验证不可靠,易产生虚假安全假象。

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA

论文配图:Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA
图 1 · 摘自论文原文
  • 分解验证器行为为判别能力与同意偏差,定位可靠性边界
  • 知识密集型任务中自验证错误率超60%,虚假接受错误答案普遍
  • 适合医疗AI安全评估者、临床决策系统开发者阅读

自验证——通过在新上下文中重新调用同一视觉语言模型(VLM)来检验自身生成的答案——正被广泛用作医学视觉问答(VQA)的安全层。本文指出该做法本质不可靠。我们提出[METHOD NAME],一种诊断框架,通过分解验证器行为的判别能力与同意偏差,映射医学VLM自验证的可靠性边界。由于验证器与生成器能力耦合,验证器会过度认同生成结果,形成‘验证幻觉’:在高验证错误率和高同意偏差并存的区域,错误答案被错误地接受。在五个医学VQA数据集和七项医学任务上评估六种开源权重的VLM,发现该边界强烈依赖任务类型:知识密集型临床任务最深陷幻觉(错误率>60%),简单任务较抵抗,感知任务居中。验证未能提供独立安全信号:逻辑混合效应分析显示,当生成器出错时,验证错误与同意偏差更可能同时发生;显著性分析揭示验证器对图像证据的关注低于生成器,称为‘懒惰验证器’。跨验证可减轻但无法消除幻觉。当验证在多轮对话中重复使用时,多数初始错误答案会被虚假验证锁定。因实验基于干净基准,实际临床部署中的失败可能更严重。

原文摘要 · Abstract (English)

Self-verification, re-invoking the same vision language model (VLM) in a fresh context to check its own generated answer, is increasingly used as a default safety layer for medical visual question answering (VQA). We argue that this practice is fundamentally unreliable. We introduce [METHOD NAME], a diagnostic framework for mapping the reliability boundary of medical VLM self-verification by decomposing verifier behavior into discrimination capability and agreement bias. Because the verifier and answer generator are capacity-coupled, the verifier can overly agree with the generator, creating a verification mirage: a regime with both high verifier error and high agreement bias, driven by false acceptance of incorrect answers. Evaluating six open-weight VLMs across five medical VQA datasets and seven medical tasks, we find that this boundary is strongly task-conditioned. Knowledge-intensive clinical tasks fall deepest into the mirage, simpler tasks are more resistant, and perceptual tasks lie in between. Verification also fails to provide an independent safety signal: logistic mixed-effects analysis shows that verifier error and agreement bias become more likely when the generator is wrong, while saliency analyses show that verifiers under-attend to image evidence relative to generators, a phenomenon we call the lazy verifier. Cross-verification reduces but does not eliminate the mirage. Moreover, when verification is reused in multi-turn actor-verifier loops, most initially wrong answers become locked in by false verification. Since our experiments use clean benchmarks, the observed reliability boundary likely underestimates failures in real clinical deployment.

医学AI自验证幻觉检测安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。