arXiv:2606.24740cs.CV2026-06中稿 · ECCV

用可学习提示增强模型,让医学图像识别更准更少误判。

BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming

论文配图:BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming
图 1 · 摘自论文原文
  • 引入混淆感知提示专家混合机制,区分正负样本。
  • 在18个数据集上实现比现有方法更高的准确率和泛化能力。
  • 适合少样本医学图像分类任务,尤其对细微差异场景有效。

视觉-语言模型(如CLIP)在自然图像领域表现出强大泛化能力,但将其应用于生物医学影像仍具挑战:全模型微调成本高,而医学数据稀缺且类别间差异细微,参数高效适配尤为关键。视觉重编程(VR)通过在输入空间注入可学习扰动提供了一种参数高效替代方案,但现有方法多关注正类提示,忽略易混淆负类,导致细粒度医学场景中预测校准不足。本文提出BioMedVR,首个基于VR的生物医学图像适配框架,通过紧凑的可学习VR模块实现预训练VLM的少样本适配。为缓解类别混淆,设计混淆最小化机制,利用大语言模型生成的混淆感知属性与混淆抑制损失,显式减少假阳性对齐。进一步提出提示专家混合结构,包含正类专家用于主类区分、负类专家用于混淆抑制,并通过自适应门控平衡二者。在18个数据集(含11个生物医学数据集和7个自然图像基准)上的大量实验表明,BioMedVR在准确率与泛化能力上均优于现有方法,有效弥合了VR与VLM在生物医学领域的鸿沟。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) such as CLIP have demonstrated strong generalization across natural-image domains. However, adapting these models to biomedical imaging is non-trivial: full-model fine-tuning is computationally expensive, while medical data are often scarce and exhibit subtle, fine-grained inter-class differences, making parameter-efficient adaptation particularly critical. Visual Reprogramming (VR) offers a parameter-efficient alternative by injecting learnable perturbations into the input space, but existing VR approaches for VLMs mainly focus on positive class prompts and overlook confusing negatives, leading to miscalibrated predictions in fine-grained medical scenarios. We present BioMedVR, the first VR-based framework for biomedical imaging, enabling few-shot adaptation of pretrained VLMs through compact learnable VR modules. To mitigate class confusion, we introduce a Confusion Minimization Mechanism that leverages LLM-generated confusion-aware attributes together with a Confusion-Suppression Loss to explicitly reduce false-positive alignment. Moreover, the designed Mixture-of-Prompt Experts combines a positive expert for main-class discrimination and a negative expert for confusion suppression, balanced via adaptive gating. Extensive experiments on 18 datasets, including 11 biomedical datasets and 7 natural image benchmarks, demonstrate that BioMedVR achieves superior accuracy and generalization, effectively bridging VR and VLMs in biomedical domains.

视觉重编程医学图像提示工程少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。