用新方法让医疗视觉语言模型在小样本下更可信。
Trustworthy Few-Shot Transfer of Medical VLMs through Split Conformal Prediction
- 提出无监督联合校准与测试数据的迁移适配策略
- 在小样本医疗图像任务中提升预测效率与覆盖率
- 适合需要可靠小样本推理的医学AI应用
医疗视觉语言模型(VLMs)展现出卓越的迁移能力,被广泛用于数据高效的图像分类。然而其可靠性尚未充分研究。本文采用分割置信区间(SCP)框架,在仅有少量标注校准数据的情况下为模型迁移提供可信保障。由于预训练通用性可能影响特定任务的置信集性质,传统微调方式会破坏SCP所需的交换性假设,导致性能下降。为此,我们提出横贯式分割置信适应(SCA-T),在不改变交换性前提下,对校准与测试数据进行无监督联合适配。在多种医学图像模态、迁移任务和非一致性评分下,实验表明该框架相比原版SCP在效率与条件覆盖上均有稳定提升,且保持相同的实证保证。
原文摘要 · Abstract (English)
Medical vision-language models (VLMs) have demonstrated unprecedented transfer capabilities and are being increasingly adopted for data-efficient image classification. Despite its growing popularity, its reliability aspect remains largely unexplored. This work explores the split conformal prediction (SCP) framework to provide trustworthiness guarantees when transferring such models based on a small labeled calibration set. Despite its potential, the generalist nature of the VLMs' pre-training could negatively affect the properties of the predicted conformal sets for specific tasks. While common practice in transfer learning for discriminative purposes involves an adaptation stage, we observe that deploying such a solution for conformal purposes is suboptimal since adapting the model using the available calibration data breaks the rigid exchangeability assumptions for test data in SCP. To address this issue, we propose transductive split conformal adaptation (SCA-T), a novel pipeline for transfer learning on conformal scenarios, which performs an unsupervised transductive adaptation jointly on calibration and test data. We present comprehensive experiments utilizing medical VLMs across various image modalities, transfer tasks, and non-conformity scores. Our framework offers consistent gains in efficiency and conditional coverage compared to SCP, maintaining the same empirical guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。