提出新型医学影像报告生成模型,支持多种影像类型且效果领先。
MicarVLMoE: A Modern Gated Cross-Aligned Vision-Language Mixture of Experts Model for Medical Image Captioning and Report Generation
- 采用分层视觉编码与门控跨模态对齐机制,提升细节捕捉与图文匹配。
- 在4类医学影像上均达当前最佳性能,尤其在肺部和病理图像上提升显著。
- 结构可解释性强,适合临床辅助诊断场景使用。
医学影像报告(MIR)旨在从放射学图像生成结构化临床描述。现有方法在细粒度特征提取、多模态对齐及跨成像类型泛化方面存在不足,通常仅针对胸部X光,依赖普通Transformer。本文提出MicarVLMoE,一种具有门控交叉对齐融合的视觉-语言混合专家模型,以解决上述问题。其架构包括:(i) 多尺度视觉编码器(MSVE),用于在不同分辨率下捕捉解剖细节;(ii) 多头双分支潜在注意力(MDLA)模块,通过潜在瓶颈表示实现视觉-语言对齐;(iii) 可调制的混合专家(MoE)解码器,实现专家自适应专精。本研究将MIR扩展至CT扫描、眼底成像、MRI及大体病理图像,在COVCTR、MMR、PGROSS和ROCO数据集上均取得最新最优结果。大量实验与消融分析验证了模型在临床准确性、跨模态对齐性及可解释性上的提升。代码已开源:https://github.com/AI-14/micar-vl-moe。
原文摘要 · Abstract (English)
Medical image reporting (MIR) aims to generate structured clinical descriptions from radiological images. Existing methods struggle with fine-grained feature extraction, multimodal alignment, and generalization across diverse imaging types, often relying on vanilla transformers and focusing primarily on chest X-rays. We propose MicarVLMoE, a vision-language mixture-of-experts model with gated cross-aligned fusion, designed to address these limitations. Our architecture includes: (i) a multiscale vision encoder (MSVE) for capturing anatomical details at varying resolutions, (ii) a multihead dual-branch latent attention (MDLA) module for vision-language alignment through latent bottleneck representations, and (iii) a modulated mixture-of-experts (MoE) decoder for adaptive expert specialization. We extend MIR to CT scans, retinal imaging, MRI scans, and gross pathology images, reporting state-of-the-art results on COVCTR, MMR, PGROSS, and ROCO datasets. Extensive experiments and ablations confirm improved clinical accuracy, cross-modal alignment, and model interpretability. Code is available at https://github.com/AI-14/micar-vl-moe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。