提升医学视觉语言模型的不确定性校准,让预测更准确高效。
LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs
- 基于图像邻近图平滑零样本概率,无需训练和标签更新。
- 显著缩小预测集合大小,降低类别间覆盖率差距。
- 黑盒轻量级设计,适合临床部署,支持无标签或少量标签使用。
医学视觉语言模型(VLMs)在医学影像中具有强大的零样本识别能力,但其在领域迁移下的可靠性依赖于可校准的不确定性。分割共形预测(SCP)提供有限样本覆盖保证,但在少样本、不平衡场景下预测集过大(效率低),且类别间覆盖率不均衡(高类别条件覆盖率差距,CCV)。直接适应校准标签会破坏交换性,导致保证失效。本文提出LATA(拉普拉斯辅助归纳适应),一种无需训练和标签的精炼方法,在联合校准与测试池上通过小规模的CCCP均场更新,沿图像-图像k-NN图平滑零样本概率,利用确定性变换保持SCP有效性。进一步引入一种故障感知共形评分,嵌入视觉语言不确定性(ViLU)框架,提供实例难度与标签合理性,提升预测集效率和类别间平衡性。LATA为黑盒、计算轻量(窗口化归纳,无需反向传播),并含可选先验控制,支持完全无标签或仅用一次校准边缘信息的标签感知变体。在三个医学VLM和九个下游任务中,LATA始终减小预测集大小并降低CCV,同时匹配或收紧目标覆盖率,优于现有归纳基线,接近标签使用方法表现,且计算开销远低于后者。全面消融与定性分析表明,LATA在不破坏交换性的前提下增强零样本预测精度。
原文摘要 · Abstract (English)
Medical vision-language models (VLMs) are strong zero-shot recognizers for medical imaging, but their reliability under domain shift hinges on calibrated uncertainty with guarantees. Split conformal prediction (SCP) offers finite-sample coverage, yet prediction sets often become large (low efficiency) and class-wise coverage unbalanced-high class-conditioned coverage gap (CCV), especially in few-shot, imbalanced regimes; moreover, naively adapting to calibration labels breaks exchangeability and voids guarantees. We propose \texttt{\textbf{LATA}} (Laplacian-Assisted Transductive Adaptation), a \textit{training- and label-free} refinement that operates on the joint calibration and test pool by smoothing zero-shot probabilities over an image-image k-NN graph using a small number of CCCP mean-field updates, preserving SCP validity via a deterministic transform. We further introduce a \textit{failure-aware} conformal score that plugs into the vision-language uncertainty (ViLU) framework, providing instance-level difficulty and label plausibility to improve prediction set efficiency and class-wise balance at fixed coverage. \texttt{\textbf{LATA}} is black-box (no VLM updates), compute-light (windowed transduction, no backprop), and includes an optional prior knob that can run strictly label-free or, if desired, in a label-informed variant using calibration marginals once. Across \textbf{three} medical VLMs and \textbf{nine} downstream tasks, \texttt{\textbf{LATA}} consistently reduces set size and CCV while matching or tightening target coverage, outperforming prior transductive baselines and narrowing the gap to label-using methods, while using far less compute. Comprehensive ablations and qualitative analyses show that \texttt{\textbf{LATA}} sharpens zero-shot predictions without compromising exchangeability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。