用视觉语言专家混合模型提升医学影像分割精度
MoME: Mixture of Visual Language Medical Experts for Medical Imaging Segmentation
- 采用多尺度视觉特征与文本嵌入动态选择专家
- 在3410例CT扫描上达到领先分割性能
- 适合需要高精度医学图像分析的研究者
本研究提出MoME,一种用于医学图像分割的视觉语言专家混合模型。该架构借鉴大语言模型中成功的混合专家(MoE)范式,通过融合多尺度视觉特征与文本嵌入,实现对医学图像复杂性的有效建模。利用涵盖3,410例CT扫描的10个数据集,在全面的医学影像分割基准上验证了其优异表现。该方法探索了基础模型在医学影像领域的集成路径,借助文本信息增强模型能力,实现了跨多个数据集的竞争力分割精度,为医学图像分析提供了新颖且稳健的架构。
原文摘要 · Abstract (English)
In this study, we propose MoME, a Mixture of Visual Language Medical Experts, for Medical Image Segmentation. MoME adapts the successful Mixture of Experts (MoE) paradigm, widely used in Large Language Models (LLMs), for medical vision-language tasks. The architecture enables dynamic expert selection by effectively utilizing multi-scale visual features tailored to the intricacies of medical imagery, enriched with textual embeddings. This work explores a novel integration of vision-language models for this domain. Utilizing an assembly of 10 datasets, encompassing 3,410 CT scans, MoME demonstrates strong performance on a comprehensive medical imaging segmentation benchmark. Our approach explores the integration of foundation models for medical imaging, benefiting from the established efficacy of MoE in boosting model performance by incorporating textual information. Demonstrating competitive precision across multiple datasets, MoME explores a novel architecture for achieving robust results in medical image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。