arXiv:2506.08356cs.CV2025-06被引 11

针对医学影像模态差异,动态分配专家模型提取特征。

MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding

  • 按报告类型路由特征,激活专用于特定影像的专家分支。
  • 在多尺度特征金字塔上实现空间自适应注意力,提升定位精度。
  • 无需额外标注,适用于多种医学影像模态,适合临床视觉语言系统。

不同医学影像模态在空间分辨率上呈现从全局粗略模式到局部细微结构的差异。然而,现有医学视觉-语言框架通常采用统一的局部特征提取策略,忽视了模态特异性需求。本文提出MedMoE,一个模块化且可扩展的视觉-语言处理框架,能根据诊断上下文动态调整视觉表征。MedMoE引入基于报告类型的Mixture-of-Experts(MoE)模块,将多尺度图像特征路由至专门训练以捕捉模态特异性视觉语义的专家分支。这些专家基于Swin Transformer主干网络生成的特征金字塔运行,支持对临床相关区域的空间自适应关注。该框架生成与文本描述对齐的局部视觉表征,且推理时无需模态特定监督。在多个医学基准上的实证结果表明,MedMoE在跨模态对齐与检索任务中均取得提升,验证了模态特异性视觉表征在临床视觉-语言系统中的价值。

原文摘要 · Abstract (English)

Different medical imaging modalities capture diagnostic information at varying spatial resolutions, from coarse global patterns to fine-grained localized structures. However, most existing vision-language frameworks in the medical domain apply a uniform strategy for local feature extraction, overlooking the modality-specific demands. In this work, we present MedMoE, a modular and extensible vision-language processing framework that dynamically adapts visual representation based on the diagnostic context. MedMoE incorporates a Mixture-of-Experts (MoE) module conditioned on the report type, which routes multi-scale image features through specialized expert branches trained to capture modality-specific visual semantics. These experts operate over feature pyramids derived from a Swin Transformer backbone, enabling spatially adaptive attention to clinically relevant regions. This framework produces localized visual representations aligned with textual descriptions, without requiring modality-specific supervision at inference. Empirical results on diverse medical benchmarks demonstrate that MedMoE improves alignment and retrieval performance across imaging modalities, underscoring the value of modality-specialized visual representations in clinical vision-language systems.

医学视觉多模态专家混合影像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。