让医学视觉语言模型学会判断不确定性的高效微调方法
Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning

- 仅更新0.11%参数,通过低维令牌更新和不确定性估计实现轻量微调
- 在15个医学影像数据集上表现优于现有方法,尤其在少样本和跨域场景
- 适合临床部署,能根据证据强弱自适应调整,提升模型鲁棒性
视觉-语言基础模型的参数高效适配对精准理解医学图像至关重要,但现有方法多为确定性且在领域偏移或图像-文本对齐模糊时表现不佳。本文提出Evi-Steer,一种面向BiomedCLIP的证据驱动跨模态低维调控框架,仅更新0.11%的总参数,同时在视觉与文本编码器中进行轻量级低维令牌更新,并估计认知不确定性。不确定性估计用于更新门控残差,使模型在证据不足时保守适应。此外,引入基于Dempster-Shafer理论的跨模态置信度融合,使视觉适应依赖于文本置信度,抑制冲突或不确定的跨模态更新。我们在涵盖8个器官、8种成像模态的15个生物医学影像数据集上进行了全面评估,覆盖少样本学习与领域泛化场景。Evi-Steer在少样本和领域偏移设置下持续优于当前最优方法,为视觉-语言模型在真实临床环境中的部署提供了实用且稳健的路径。代码已开源。
原文摘要 · Abstract (English)
Parameter-efficient adaptation of vision-language foundation models is crucial for precise multimodal understanding of biomedical images, yet existing methods remain deterministic and often struggle under domain shift or ambiguous image-text alignment. This limitation is particularly critical in the clinic, where models should remain robust in low-data regimes and domain shifts. We present Evi-Steer, an evidential cross-modal low-dimensional steering framework for BiomedCLIP that enables uncertainty-aware parameter-efficient fine-tuning while updating only 0.11% of total model parameters. Our approach performs lightweight low-dimensional token updates in both vision and text encoders while simultaneously estimating epistemic uncertainty. These uncertainty estimates update gate residuals, allowing the model to adapt conservatively when evidence is weak. Furthermore, we introduce cross-modal confidence fusion based on Dempster-Shafer theory, enabling visual adaptation to be conditioned on textual confidence and suppressing conflicting or uncertain cross-modal updates. We conduct a comprehensive evaluation on 15 biomedical imaging datasets spanning 8 organs and 8 imaging modalities under few-shot learning and domain generalization settings. Evi-Steer consistently outperforms state-of-the-art methods under few-shot learning and domain shift settings, demonstrating a practical and robust pathway for deploying vision-language models in real-world clinical settings. Code is available at https://github.com/HealthX-Lab/Evi-Steer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。