arXiv:2511.09540cs.CV2025-11AAAI被引 2

提出在球面流形上统一对齐医学视觉语言模型的提示,提升少样本分类性能。

vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs

  • 在统一球面流形上建模提示分布,实现跨模型语义对齐
  • 在14个医疗数据集上准确率显著优于现有方法
  • 适合需要少样本适应的医学影像分析场景

基于大语言模型(LLM)引导的上下文优化(CoOp)为生物医学CLIP类视觉语言模型(VLMs)提供了可扩展的提示工程替代方案。然而,由于训练语料和模型架构差异,LLM与CLIP变体之间存在语义错位;且传统欧氏空间优化难以建模统一表征或施加局部几何约束,加剧了复杂医学影像中的模态差距并导致少样本适应不稳定。本文提出vMFCoOp,通过在共享球面流形上逆向估计冯·米塞斯-费舍尔(vMF)分布,利用统一语义锚点对齐任意LLM与CLIP主干的语义偏移,实现稳健的生物医学提示与更优的少样本分类。该框架基于三项互补约束,在14个医学数据集、12种医学影像模态和13个解剖区域上表现一致提升,超越现有最先进方法,在准确率、泛化性和临床适用性方面均具优势。本工作旨在持续拓展下游应用,相关资源将通过https://github.com/VinyehShaw/UniEqui 公开共享。

原文摘要 · Abstract (English)

Recent advances in context optimization (CoOp) guided by large language model (LLM)-distilled medical semantic priors offer a scalable alternative to manual prompt engineering and full fine-tuning for adapting biomedical CLIP-based vision-language models (VLMs). However, prompt learning in this context is challenged by semantic misalignment between LLMs and CLIP variants due to divergent training corpora and model architectures; it further lacks scalability across continuously evolving families of foundation models. More critically, pairwise multimodal alignment via conventional Euclidean-space optimization lacks the capacity to model unified representations or apply localized geometric constraints, which tends to amplify modality gaps in complex biomedical imaging and destabilize few-shot adaptation. In this work, we propose vMFCoOp, a framework that inversely estimates von Mises-Fisher (vMF) distributions on a shared Hyperspherical Manifold, aligning semantic biases between arbitrary LLMs and CLIP backbones via Unified Semantic Anchors to achieve robust biomedical prompting and superior few-shot classification. Grounded in three complementary constraints, vMFCoOp demonstrates consistent improvements across 14 medical datasets, 12 medical imaging modalities, and 13 anatomical regions, outperforming state-of-the-art methods in accuracy, generalization, and clinical applicability. This work aims to continuously expand to encompass more downstream applications, and the corresponding resources are intended to be shared through https://github.com/VinyehShaw/UniEqui.

视觉语言模型医学影像提示学习球面优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。