arXiv:2503.05626cs.CVcs.AI2025-03被引 12

提出FMT模型,提升肺炎多模态诊断在数据缺失下的准确率。

FMT:A Multimodal Pneumonia Detection Model Based on Stacking MOE Framework

  • 用ResNet-50与BERT联合建模,动态模拟模态缺失增强鲁棒性。
  • 在小规模数据集上达到94%准确率、95%召回率,优于现有方法。
  • 适合资源受限场景,可扩展至其他医学多模态任务。

人工智能在肺炎诊断的医学图像分析中展现出提升准确性的潜力。然而,传统多模态方法难以应对真实世界中的数据不完整和模态缺失问题。本研究提出灵活多模态变换器(FMT),采用ResNet-50与BERT进行联合表示学习,并引入动态掩码注意力机制模拟临床模态缺失,以增强模型鲁棒性;最后使用序列式专家混合(MOE)架构实现多层次决策优化。在小型多模态肺炎数据集上的评估显示,FMT达到94%准确率、95%召回率和93% F1分数,显著优于单模态基线(ResNet:89%;BERT:79%)及医疗基准CheXMed(90%),为资源有限的医疗环境提供了可扩展的肺炎多模态诊断解决方案。

原文摘要 · Abstract (English)

Artificial intelligence has shown the potential to improve diagnostic accuracy through medical image analysis for pneumonia diagnosis. However, traditional multimodal approaches often fail to address real-world challenges such as incomplete data and modality loss. In this study, a Flexible Multimodal Transformer (FMT) was proposed, which uses ResNet-50 and BERT for joint representation learning, followed by a dynamic masked attention strategy that simulates clinical modality loss to improve robustness; finally, a sequential mixture of experts (MOE) architecture was used to achieve multi-level decision refinement. After evaluation on a small multimodal pneumonia dataset, FMT achieved state-of-the-art performance with 94% accuracy, 95% recall, and 93% F1 score, outperforming single-modal baselines (ResNet: 89%; BERT: 79%) and the medical benchmark CheXMed (90%), providing a scalable solution for multimodal diagnosis of pneumonia in resource-constrained medical settings.

肺炎检测多模态MOE鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。