arXiv:2506.06076cs.CV2025-06被引 8

提升医学视觉语言模型的可靠性,让预测更可信且高效。

Full Conformal Adaptation of Medical Vision-Language Models

  • 提出全置信度自适应框架,边适配边校准,实现逐样本动态调整。
  • 在9个任务上提升27%的集合效率,同时保证预测覆盖率不变。
  • 无需训练的线性探测器降低计算开销,适合医疗场景快速部署。

大规模预训练的视觉语言模型(VLMs)展现出强大的迁移能力,正逐步融入医学图像分析。尽管其判别性能已被广泛研究,但可靠性仍被忽视。本文研究其在分裂置信预测(SCP)框架下的表现,该框架通过标定集理论上保证给定误差水平。然而,VLM的零样本性能受限,常规少样本微调无法满足SCP的交换性假设。为此,我们提出全置信度自适应,一种联合适配与校准的新型设置,通过少量适配集对每个测试样本进行归纳推理。此外,我们引入无需训练的线性探测器SS-Text,缓解此类方法的计算负担。我们在3个模态专用医学VLM和9个适配任务上进行了全面实验。本框架使用与SCP完全相同的数据,相对效率提升最高达27%,且保持相同覆盖保证。

原文摘要 · Abstract (English)

Vision-language models (VLMs) pre-trained at large scale have shown unprecedented transferability capabilities and are being progressively integrated into medical image analysis. Although its discriminative potential has been widely explored, its reliability aspect remains overlooked. This work investigates their behavior under the increasingly popular split conformal prediction (SCP) framework, which theoretically guarantees a given error level on output sets by leveraging a labeled calibration set. However, the zero-shot performance of VLMs is inherently limited, and common practice involves few-shot transfer learning pipelines, which cannot absorb the rigid exchangeability assumptions of SCP. To alleviate this issue, we propose full conformal adaptation, a novel setting for jointly adapting and conformalizing pre-trained foundation models, which operates transductively over each test data point using a few-shot adaptation set. Moreover, we complement this framework with SS-Text, a novel training-free linear probe solver for VLMs that alleviates the computational cost of such a transductive approach. We provide comprehensive experiments using 3 different modality-specialized medical VLMs and 9 adaptation tasks. Our framework requires exactly the same data as SCP, and provides consistent relative improvements of up to 27% on set efficiency while maintaining the same coverage guarantees.

医学视觉置信度校准少样本学习VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。