arXiv:2602.21154cs.AI2026-02中稿 · ICASSP 2026

通过时空掩码与表征解耦,提升心电图与病历联合建模精度

CG-DMER: Hybrid Contrastive-Generative Framework for Disentangled Multimodal ECG Representation Learning

  • 设计时空掩码重建机制,捕捉心电图多导联时序依赖关系
  • 引入解耦对齐策略,在三个数据集上实现领先性能
  • 适合心血管疾病智能诊断、多模态医疗表征学习研究者

准确解读心电图信号对诊断心血管疾病至关重要。近年来,将心电图与临床报告结合的多模态方法展现出巨大潜力,但仍存在两大模态层面问题:(1) 模态内:现有模型以导联无关方式处理心电图,忽视导联间的时空依赖性,限制了对精细诊断模式的建模能力;(2) 模态间:现有方法直接对齐心电图与自由文本报告,因报告的非结构化特性引入模态特异性偏差。针对上述问题,本文提出CG-DMER——一种基于对比-生成框架的解耦多模态心电图表示学习方法,包含两项核心设计:(1) 时空掩码建模:在空间与时间维度同时施加掩码并重建缺失信息,以更好地捕捉细粒度时序动态与导联间空间依赖;(2) 表征解耦与对齐策略:通过引入模态特异性与共享编码器,削弱冗余噪声与模态偏差,实现模态不变与模态特异性表征的清晰分离。在三个公开数据集上的实验表明,CG-DMER在多种下游任务中均达到最先进水平。

原文摘要 · Abstract (English)

Accurate interpretation of electrocardiogram (ECG) signals is crucial for diagnosing cardiovascular diseases. Recent multimodal approaches that integrate ECGs with accompanying clinical reports show strong potential, but they still face two main concerns from a modality perspective: (1) intra-modality: existing models process ECGs in a lead-agnostic manner, overlooking spatial-temporal dependencies across leads, which restricts their effectiveness in modeling fine-grained diagnostic patterns; (2) inter-modality: existing methods directly align ECG signals with clinical reports, introducing modality-specific biases due to the free-text nature of the reports. In light of these two issues, we propose CG-DMER, a contrastive-generative framework for disentangled multimodal ECG representation learning, powered by two key designs: (1) Spatial-temporal masked modeling is designed to better capture fine-grained temporal dynamics and inter-lead spatial dependencies by applying masking across both spatial and temporal dimensions and reconstructing the missing information. (2) A representation disentanglement and alignment strategy is designed to mitigate unnecessary noise and modality-specific biases by introducing modality-specific and modality-shared encoders, ensuring a clearer separation between modality-invariant and modality-specific representations. Experiments on three public datasets demonstrate that CG-DMER achieves state-of-the-art performance across diverse downstream tasks.

心电图分析多模态学习表征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。