arXiv:2508.02557eess.IVcs.CV2025-08

用强化学习动态对齐多模态心脏影像,提升3D全心分割精度。

RL-U$^2$Net: A Dual-Branch UNet with Reinforcement Learning-Assisted Multimodal Feature Fusion for Accurate 3D Whole-Heart Segmentation

  • 双分支U-Net并行处理CT与MRI,用强化学习自动调整对齐策略。
  • 在MM-WHS 2017数据集上,CT分割Dice达93.1%,MRI达87.0%。
  • 适合心血管影像分析、医学图像融合研究者参考。

准确的全心分割对心血管疾病精确诊断与介入规划至关重要。融合CT与MRI等多模态互补信息可显著提升分割精度与鲁棒性。然而现有方法存在模态间空间不一致、融合策略静态且缺乏适应性、特征对齐与分割过程脱节等问题。为此,我们提出基于强化学习辅助特征对齐的双分支U-Net模型(RL-U²Net),用于精确高效的多模态3D全心分割。该模型采用双分支U型网络并行处理CT与MRI图像块,引入新型RL-XAlign模块连接编码器。该模块通过跨模态注意力捕捉语义对应关系,并利用强化学习代理学习最优旋转策略,持续对齐解剖姿态与纹理特征。对齐后的特征经各自解码器重建,最终通过集成学习决策模块融合各块预测结果,生成整体分割图。在公开的MM-WHS 2017数据集上的实验表明,所提方法优于现有最先进方法,在CT和MRI上的Dice系数分别达到93.1%和87.0%,验证了其有效性与优越性。

原文摘要 · Abstract (English)

Accurate whole-heart segmentation is a critical component in the precise diagnosis and interventional planning of cardiovascular diseases. Integrating complementary information from modalities such as computed tomography (CT) and magnetic resonance imaging (MRI) can significantly enhance segmentation accuracy and robustness. However, existing multi-modal segmentation methods face several limitations: severe spatial inconsistency between modalities hinders effective feature fusion; fusion strategies are often static and lack adaptability; and the processes of feature alignment and segmentation are decoupled and inefficient. To address these challenges, we propose a dual-branch U-Net architecture enhanced by reinforcement learning for feature alignment, termed RL-U$^2$Net, designed for precise and efficient multi-modal 3D whole-heart segmentation. The model employs a dual-branch U-shaped network to process CT and MRI patches in parallel, and introduces a novel RL-XAlign module between the encoders. The module employs a cross-modal attention mechanism to capture semantic correspondences between modalities and a reinforcement-learning agent learns an optimal rotation strategy that consistently aligns anatomical pose and texture features. The aligned features are then reconstructed through their respective decoders. Finally, an ensemble-learning-based decision module integrates the predictions from individual patches to produce the final segmentation result. Experimental results on the publicly available MM-WHS 2017 dataset demonstrate that the proposed RL-U$^2$Net outperforms existing state-of-the-art methods, achieving Dice coefficients of 93.1% on CT and 87.0% on MRI, thereby validating the effectiveness and superiority of the proposed approach.

医学影像多模态分割强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。