arXiv:2606.16583cs.CL2026-06

现有不确定性估计无法保障临床模型安全,但可预警模型失效。

Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?

  • 用扰动测试发现模型在关键情况下的不确定性几乎不变
  • 未扰动输入的不确定性能提前预测模型崩溃点
  • 建议用扰动评估替代传统校准,适合临床部署前检验

临床视觉问答(VQA)模型的安全部署依赖可靠的不确定性估计(UE),即判断预测是否可信的信号。我们对12个临床VLMs上的8种UE方法进行基准测试,发现UE质量并非方法固有属性:其表现随模型准确率下降而恶化,恰在最需可靠性的薄弱环节失效。当通过隐藏正确选项(NOTA扰动)施压模型时,准确率大幅下降,但不确定性几乎不变,导致系统性失准。然而,未扰动输入的不确定性可有效预判哪些预测将在扰动下崩溃,表明当前VLM的不确定性蕴含模型脆弱性的诊断信息。研究将UE定位为识别脆弱预测的诊断工具,并提出基于扰动的评估是实现临床安全部署的关键路径。

原文摘要 · Abstract (English)

Safe deployment of clinical vision-language models (VLMs) requires reliable uncertainty estimation (UE): a signal indicating when predictions should be trusted or escalated to a clinician. We test whether current UE methods actually deliver this signal. Benchmarking 8 methods across 12 VLMs on clinical visual question-answering (VQA), we find that UE quality is not an intrinsic property of the UE method: it tracks model accuracy, degrading precisely where the model performance is weakest, and therefore where reliability is most needed. When we stress-test models by hiding the correct option among the multiple-choice answers (NOTA perturbations), accuracy collapses while uncertainty barely changes, leaving models systematically miscalibrated. Yet, we find that uncertainty on the unperturbed input reliably anticipates which predictions will collapse under NOTA, indicating that UE in current VLMs carries diagnostic information about model fragility. Our results position UE as a diagnostic tool for identifying fragile predictions and motivate perturbation-based evaluation as a path toward safe clinical deployment.

临床AI不确定性估计模型鲁棒性视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。