arXiv:2606.25535cs.CV2026-06

用不完整MRI数据生成精准肿瘤定量图,突破现有方法局限。

Spatio-Temporal Mixture-of-Modality-Experts Diffusion for Quantitative DCE-MRI Synthesis from Incomplete MR Sequences

论文配图:Spatio-Temporal Mixture-of-Modality-Experts Diffusion for Quantitative DCE-MRI Synthesis from Incomplete MR Sequences
图 1 · 摘自论文原文
  • 设计时空门控专家网络,动态融合多模态影像特征
  • 在386例脑瘤患者上实现最低整体误差,肿瘤区重建最准
  • 适合临床缺失数据场景,可解释的融合机制贴近医学逻辑

动态对比增强MRI(DCE-MRI)的定量参数对肿瘤评估至关重要,但常因造影剂风险和扫描协议差异无法获取。已有方法依赖固定、完整的输入模态,难以应对真实数据缺失。本文提出时空混合模态专家扩散模型(ST-MoME),从多样化的多模态MRI子集合成3D DCE参数图。该模型通过时空门控网络融合各模态专家特征,生成体素级、时间步依赖的权重,构成条件张量引导去噪过程。为保障定量精度,采用3D块状训练与Swin骨干网络,在图像空间直接进行扩散建模。在386例临床脑瘤患者数据上,评估了16种受控模态可用性场景。结果表明,其在所有三类DCE参数上的平均归一化均方误差(NMSE)最低,对$ v_p $和$ v_e $表现领先,$ K^{ ext{trans}} $具竞争力,且在临床关键肿瘤区域误差最小。后验分析显示,学习到的门控动态呈现结构早、生理晚的融合顺序,符合临床直觉。

原文摘要 · Abstract (English)

Quantitative maps from dynamic contrast-enhanced MRI (DCE-MRI) are essential for tumor assessment but are often unavailable due to contrast-agent risks and protocol variability. Prior methods predict these maps from other MRI modalities, yet most assume fixed, fully observed inputs and fail under realistic missingness. We present Spatio-Temporal Mixture-of-Modality-Experts (ST-MoME), a conditional diffusion framework that synthesizes 3D DCE parameter maps from diverse subsets of multimodal MRI. ST-MoME fuses modality-specific expert features through a spatio-temporal gating network that produces voxel-wise, timestep-dependent weights, forming a conditioning tensor that guides denoising. To preserve quantitative fidelity, ST-MoME performs diffusion directly in image space with 3D patch-based training and a Swin-based backbone. On a clinical brain-tumor cohort of 386 patients, we evaluate ST-MoME across 16 controlled modality-availability scenarios. It achieves the lowest mean Normalized Mean Square Error (NMSE) aggregated across all three DCE parameters, with leading performance on $v_p$ and $v_e$, competitive results on $K^{\mathrm{trans}}$, and the lowest reconstruction error within the clinically critical tumor region. A post-hoc analysis of the learned gating dynamics shows a structural-early, physiological-late fusion schedule consistent with clinical intuition.

医学影像扩散模型多模态融合定量成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。