医学视觉语言模型需靠可信置信度判断何时该暂停,而非自主决策。
Calibrated Triage, Not Autonomy: Confidence Estimation for Medical Vision-Language Models
- 评估九种置信度估计方法在医疗影像上的表现
- 仅约四分之一放射科病例可安全自动转诊,病理几乎不可靠
- 置信度应作为临床监督下的分诊工具,非自主判断依据
视觉语言模型能流畅回答胸部X光或病理切片相关问题,却可能严重依赖语言先验而忽略图像。在医学中,这种看似可信实则错误的输出危害最大,因此可靠的置信度评分是关键防线。本文不追求准确率,而是关注模型在何种置信信号下可安全自主处理多少医疗影像工作。我们在五种开源多模态模型和三个医学VQA数据集上评估了九种置信度估计器,涵盖无训练的对数基线、基于提示的自报告及训练好的内部探针,所有探针均仅用自然图像训练且未适配医学数据。结果表明:标准指标具有误导性,各估计器间区分度极低;固定高置信阈值的差异被夸大,因分数尺度不可比。在20%误差容忍下,最强估计器可在分布无关保证下使约25%放射科案例安全转诊,持留阈值下达三分之一,但病理任务几乎无法实现。可靠角色是经临床监督的校准分诊,而非自主决策——好估计器能让有能力的模型安全延迟,但无法凭空制造可靠性。
原文摘要 · Abstract (English)
A vision-language model can answer a question about a chest radiograph or a pathology slide fluently and confidently while barely using the image, relying instead on language priors. In medicine this is the failure that matters most: the answer looks trustworthy and is not, and the natural safeguard is a confidence score reliable enough to say when the model should abstain. We ask a deployment question rather than an accuracy one: how much medical imaging work a vision-language model can safely defer on its own, and which confidence signal makes that possible. We evaluate nine confidence estimators, spanning training-free logit baselines, prompt-based self-reports, and trained internal probes, across five open-weight LVLMs and three medical VQA datasets covering broad clinical imaging, radiology, and pathology, every probe trained only on natural images and applied to medicine without adaptation. Recast as bounded selective prediction, the comparison is cautionary. Standard metrics mislead: discrimination barely separates the estimators, and a fixed high-confidence cutoff separates them far less than it appears, because their scores sit on incomparable scales; no estimator is reliably best across domains or models. What can be safely deferred is set at two levels: base-model competence fixes a ceiling, and the confidence layer determines how much of it is reachable. At a 20% error tolerance the strongest estimator defers about a quarter of radiology cases under a distribution-free guarantee and a third under a held-out threshold, and little to none of pathology. The usable role is calibrated triage under clinical oversight, not autonomous deferral: a good estimator makes a competent model defer safely where it is competent, but none manufactures reliability where the base model lacks it. We release all outputs, correctness judgments, and confidence scores, with code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。