融合MRI影像与91个影像组学特征,提升脑肿瘤分类准确率
Multimodal Brain Tumour Classification Using Feature Fusion
- 双分支网络分别处理MRI图像和91个影像组学特征
- 门控融合策略达到96.13%准确率,优于单一模态模型
- 适合医学影像分析与多模态深度学习研究者参考
临床医生通过整合患者症状、病史及MRI、CT等多模态影像数据做出脑肿瘤诊断。然而,多数深度学习模型仅依赖MRI/CT图像,未能复现临床的多模态推理。本文构建双分支网络,结合原始MRI扫描与91个提取的影像组学特征(包括强度、纹理、形状和边界描述符),将脑肿瘤分为胶质瘤、脑膜瘤、垂体瘤和无肿瘤四类。预训练的CNN骨干网络处理图像流,专用MLP处理影像组学流。两种流通过拼接、门控或双向跨模态注意力策略融合。在7,200张图像的平衡数据集上进行九次实验,所有多模态配置均优于单模态基线,其中门控融合取得最高准确率96.13%。
原文摘要 · Abstract (English)
Clinicians diagnose brain tumors by synthesizing patient symptoms, medical history, and quantitative imaging data from modalities such as MRI and CT scans into a unified clinical judgement. However, most deep learning models rely on MRI/CT images alone, failing to replicate the clinicians multimodal reasoning. We explore a two-branch multimodal network combining raw MRI scans with 91 extracted radiomic features (intensity, texture, shape, and boundary descriptors) to classify brain tumors into glioma, meningioma, pituitary, and no-tumor. A pre-trained CNN backbone encodes the image stream, whereas a dedicated MLP encodes the radiomic stream. Both streams are fused via concatenation, gated, or bidirectional cross-modal attention strategies. Across nine experimental runs on a balanced 7,200 image dataset, all multimodal configurations outperform unimodal baselines with gated fusion achieving the best accuracy of 96.13%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。