分离心脏影像的视角与疾病特征,提升罕见病诊断准确性。
Motion-Guided Causal Disentanglement for Robust Multi-View Cine Cardiac MRI Diagnosis

- 用双分支对比学习和对抗约束解耦视角与疾病信息。
- 在小样本下仍保持高精度,分类准确率提升5.2%以上。
- 适合医疗影像分析、罕见病诊断等需要泛化能力的场景。
多视角心脏磁共振(CMR)成像可提供互补的解剖信息,广泛用于无创疾病评估。近年来基于Transformer的模型在CMR分析中展现出强大的表征学习能力,但通常学习统一的潜在嵌入,将视角特异的解剖变化与疾病相关特征纠缠在一起,导致分类器偏向结构属性而非视图无关的病理模式。这一问题在低数据场景下尤为严重,尤其是对未充分代表的心脏疾病,样本有限易引发捷径学习和依赖视角的决策边界。为此,我们提出基于ViT-MAE骨干网络的运动引导视图-疾病解耦框架MoViD。该模型通过双分支监督对比目标和梯度反转对抗约束,显式分解潜在表示为视角特异与疾病判别成分,并最小化疾病信息向视角嵌入泄露。此外,引入无需标注的时序运动特征(基于帧间差分图),用于定位搏动心脏区域并抑制背景伪影。对比损失中融入焦点重加权机制以缓解类别不平衡。我们在一个私有临床静脉血栓数据集及两个公开基准(M&Ms, M&Ms2)上进行评估。在疾病分类与心脏分割任务中,本方法持续优于标准Transformer基线,并在性能上媲美大规模预训练基础模型,验证了结构解耦在医学图像分析中的有效性。
原文摘要 · Abstract (English)
Multi-view cardiac magnetic resonance (CMR) imaging provides complementary anatomical information and is widely used for noninvasive disease assessment. Recent transformer-based models have demonstrated strong representation learning capabilities for CMR analysis; however, they typically learn unified latent embeddings that entangle view-specific anatomical variations with disease-related features. Such entanglement biases classifiers toward structural attributes rather than view-invariant pathological patterns. This issue is exacerbated in low-data regimes, particularly for underrepresented cardiac conditions, where limited samples increase the susceptibility to shortcut learning and view-dependent decision boundaries. To address this, we propose a Motion-Guided View--Disease Disentanglement framework MoViD built upon a ViT-MAE backbone. The model explicitly factorizes latent representations into view-specific and disease-discriminative components using dual-branch supervised contrastive objectives and a gradient-reversal adversarial constraint that minimizes disease leakage into the view embedding. Additionally, an annotation-free temporal motion feature, derived from inter-frame difference maps, is introduced to localize the beating heart region and suppress background artifacts. A focal reweighting mechanism is incorporated into the contrastive loss to mitigate class imbalance. We evaluate the framework on a private clinical venous thrombosis dataset and two public benchmarks (M&Ms, M&Ms2). Across disease classification and cardiac segmentation tasks, our approach consistently outperforms standard transformer baselines and demonstrates competitive performance against large-scale pretrained foundation models, validating the efficacy of structural disentanglement in medical image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。