融合视觉语义空间信息,提升脑肿瘤分割精度与边界稳定性。
Unified Multimodal Coherent Field: Synchronous Semantic-Spatial-Vision Fusion for Brain Tumor Segmentation
- 在统一3D潜空间中同步融合视觉、语义与空间信息,自适应调整模态贡献。
- 在BraTS 2020/2021上平均Dice达0.8579和0.8977,性能提升4.18%。
- 结合医学先验知识设计注意力机制,适合临床高精度分割场景。
脑肿瘤分割需从多序列磁共振成像(MRI)中准确识别全肿瘤(WT)、肿瘤核心(TC)和增强肿瘤(ET)等层次区域。由于肿瘤组织异质性、边界模糊及多序列间对比度差异,仅依赖视觉信息或事后损失约束的方法在边界勾画与层次保持上表现不稳定。为此,本文提出统一多模态相干场(UMCF)方法,在统一的3D潜空间中实现视觉、语义与空间信息的同步交互融合,通过无参数不确定性门控自适应调节模态权重,并将医学先验知识直接嵌入注意力计算,避免传统“处理-拼接”分离架构。在BraTS 2020与2021数据集上,UMCF+nnU-Net分别取得0.8579与0.8977的平均Dice系数,较主流架构平均提升4.18%。通过深度整合临床知识与影像特征,为精准医疗中的多模态信息融合提供新路径。
原文摘要 · Abstract (English)
Brain tumor segmentation requires accurate identification of hierarchical regions including whole tumor (WT), tumor core (TC), and enhancing tumor (ET) from multi-sequence magnetic resonance imaging (MRI) images. Due to tumor tissue heterogeneity, ambiguous boundaries, and contrast variations across MRI sequences, methods relying solely on visual information or post-hoc loss constraints show unstable performance in boundary delineation and hierarchy preservation. To address this challenge, we propose the Unified Multimodal Coherent Field (UMCF) method. This method achieves synchronous interactive fusion of visual, semantic, and spatial information within a unified 3D latent space, adaptively adjusting modal contributions through parameter-free uncertainty gating, with medical prior knowledge directly participating in attention computation, avoiding the traditional "process-then-concatenate" separated architecture. On Brain Tumor Segmentation (BraTS) 2020 and 2021 datasets, UMCF+nnU-Net achieves average Dice coefficients of 0.8579 and 0.8977 respectively, with an average 4.18% improvement across mainstream architectures. By deeply integrating clinical knowledge with imaging features, UMCF provides a new technical pathway for multimodal information fusion in precision medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。