arXiv:2509.05154eess.IVcs.CV2025-09被引 1

用简单CNN集成多个视觉语言模型,显著提升医学图像分割精度。

VLSM-Ensemble: Ensembling CLIP-based Vision-Language Models for Enhanced Medical Image Segmentation

  • 用低复杂度CNN融合多个基于CLIP的视觉语言模型进行分割
  • 在BKAI肠镜数据集上Dice分数提升6.3%,其他数据集增1%-6%
  • 方法适用于多种医学影像,为未来研究提供新方向

视觉语言模型及其在图像分割任务中的应用展现出生成高精度、可解释结果的巨大潜力。然而,基于CLIP和BiomedCLIP的实现仍落后于更复杂的架构如CRIS。本文不关注文本提示工程,而是通过引入一个低复杂度的CNN来集成视觉语言分割模型(VLSM),有效缩小差距。实验显示,使用集成的BiomedCLIPSeg在BKAI肠镜数据集上实现了6.3%的Dice分数提升,其他数据集增益为1%至6%。此外,我们在四个放射科与非放射科数据集上获得初步结果。结果显示,集成效果在不同数据集上表现各异(从优于到弱于CRIS模型),表明该现象值得社区深入探究。代码已开源:https://github.com/juliadietlmeier/VLSM-Ensemble。

原文摘要 · Abstract (English)

Vision-language models and their adaptations to image segmentation tasks present enormous potential for producing highly accurate and interpretable results. However, implementations based on CLIP and BiomedCLIP are still lagging behind more sophisticated architectures such as CRIS. In this work, instead of focusing on text prompt engineering as is the norm, we attempt to narrow this gap by showing how to ensemble vision-language segmentation models (VLSMs) with a low-complexity CNN. By doing so, we achieve a significant Dice score improvement of 6.3% on the BKAI polyp dataset using the ensembled BiomedCLIPSeg, while other datasets exhibit gains ranging from 1% to 6%. Furthermore, we provide initial results on additional four radiology and non-radiology datasets. We conclude that ensembling works differently across these datasets (from outperforming to underperforming the CRIS model), indicating a topic for future investigation by the community. The code is available at https://github.com/juliadietlmeier/VLSM-Ensemble.

医学图像视觉语言模型集成分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。