用诊断意图指导医学影像融合,提升肿瘤分割精度。
MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion

- 基于扩散变换器与诊断意图文本生成融合策略。
- 在3个数据集上显著提升肿瘤分割准确率。
- 适合需要交互式智能辅助诊断的临床场景。
医学图像融合旨在整合不同成像模态的互补信息以支持临床诊断。现有方法通常采用全局统一的融合规则,缺乏对诊断意图和病灶结构的深入理解。为此,我们提出MIND,一种基于扩散变换器(DiTs)的多模态意图驱动网络。具体地,利用BioMedGPT从源图像生成意图驱动的融合文本,以病理感知的诊断意图引导融合过程。为解决DiTs中1D序列展平导致的2D空间连续性损失,设计了多尺度潜在适配器模块,该模块在序列化前显式提取源图像特征,并通过严格维度对齐注入网络,有效补充图像特征。为缓解图像输出与诊断意图解耦引发的语义漂移,设计医学语义一致性损失,确保融合图像与融合文本之间的深层语义锁定,同时保持底层物理流形重建的稳定性。在Harvard、BraTS和GFP数据集上的综合实验表明,MIND实现了更优的融合质量,显著提升了下游脑肿瘤分割精度,并支持灵活的交互式融合,为意图驱动的智能临床决策支持系统提供了重要前景。
原文摘要 · Abstract (English)
Medical image fusion aims to integrate complementary information from diverse imaging modalities to support clinical diagnosis. Existing methods typically apply uniform fusion rules globally, lacking a deep understanding of diagnostic intents and pathological structures. To address these limitations, we propose MIND, a Multimodal Intent-Driven Network via Diffusion Transformers (DiTs) for medical image fusion. Specifically, we utilize BioMedGPT to generate intent-driven fusion texts from source images, guiding the fusion process with pathology-aware diagnostic intents. To combat the loss of 2D spatial continuity caused by 1D sequence flattening in DiTs, we design a Multi-scale Latent Adapter. This module explicitly extracts source image features before serialization, injecting them into the network via strict dimensional alignment to effectively supplement image features. To resolve the semantic shift caused by decoupling image outputs from diagnostic intents, we design a medical semantic consistency loss. This loss ensures deep semantic locking between fused images and fusion texts while maintaining the stability of the underlying physical manifold reconstruction. Comprehensive experiments on the Harvard, BraTS, and GFP datasets reveal that MIND delivers superior fusion quality, significantly improves downstream brain tumor segmentation accuracy, and enables flexible interactive fusion, holding significant promise for intent-driven intelligent clinical decision support systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。