用多模态物理生成与视觉语言引导,提升脑肿瘤分割精度。
Multimodal Fusion at Three Tiers: Physics-Driven Data Generation and Vision-Language Guidance for Brain Tumor Segmentation
- 分像素、特征、语义三层融合,逐级整合多模态信息。
- 在BraTS数据集上平均Dice达0.89,HD95降低6.57毫米。
- 适合医学影像分析与跨模态学习研究者参考。
精准的脑肿瘤分割对神经肿瘤诊断与治疗规划至关重要。深度学习虽有进展,但自动分割仍面临肿瘤形态异质性和复杂三维空间关系等挑战。本文提出一种三层融合架构,实现精确脑肿瘤分割。在像素层,通过物理建模将磁共振成像(MRI)扩展为包含模拟超声和合成计算机断层扫描(CT)的多模态数据;在特征层,采用基于Transformer的跨模态特征融合,通过多教师协同蒸馏整合三种专家模型(MRI、US、CT);在语义层,利用GPT-4V生成临床文本知识,经CLIP对比学习与特征逐通道调制(FiLM)转化为空间引导信号。三者构成从数据增强到特征提取再到语义引导的完整流程。在BraTS 2020、2021、2023数据集上验证,平均Dice系数分别为0.8665、0.9014、0.8912,较基线平均降低95%豪斯多夫距离(HD95)6.57毫米。该方法为精准肿瘤分割与边界定位提供新范式。
原文摘要 · Abstract (English)
Accurate brain tumor segmentation is crucial for neuro-oncology diagnosis and treatment planning. Deep learning methods have made significant progress, but automatic segmentation still faces challenges, including tumor morphological heterogeneity and complex three-dimensional spatial relationships. This paper proposes a three-tier fusion architecture that achieves precise brain tumor segmentation. The method processes information progressively at the pixel, feature, and semantic levels. At the pixel level, physical modeling extends magnetic resonance imaging (MRI) to multimodal data, including simulated ultrasound and synthetic computed tomography (CT). At the feature level, the method performs Transformer-based cross-modal feature fusion through multi-teacher collaborative distillation, integrating three expert teachers (MRI, US, CT). At the semantic level, clinical textual knowledge generated by GPT-4V is transformed into spatial guidance signals using CLIP contrastive learning and Feature-wise Linear Modulation (FiLM). These three tiers together form a complete processing chain from data augmentation to feature extraction to semantic guidance. We validated the method on the Brain Tumor Segmentation (BraTS) 2020, 2021, and 2023 datasets. The model achieves average Dice coefficients of 0.8665, 0.9014, and 0.8912 on the three datasets, respectively, and reduces the 95% Hausdorff Distance (HD95) by an average of 6.57 millimeters compared with the baseline. This method provides a new paradigm for precise tumor segmentation and boundary localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。