用信息论方法学习心电图的多模态表示,提升细粒度诊断性能。
Information-theoretic Multimodal Representation Learning for Electrocardiogram Signals

- 基于信息论构建联合优化目标,同时保留心电波形结构与临床语义。
- 在PTB-XL数据集上,分类F1提升超3%,子类分类提升超5%。
- 适用于需要精细心电分析的临床场景,如疾病亚型识别。
心电图(ECG)是心脏活动的无创测量,在临床诊断中起核心作用。现有方法将心电图信号与临床报告对齐以引入诊断语义,但临床报告常无法保留心电波形丰富的生理结构,尤其在从粗略诊断类别到精细形态学的不同抽象层次上。为此,本文从信息论角度出发,推导出一个可计算的目标函数,兼顾信号结构保留与临床语义融合。基于此,提出 extbf{MERIT}(Multimodal ECG Representation via Information Theory),一种结合掩码心电图建模与心电图-文本对比对齐的双分支预训练框架。在PTB-XL及其他基准上的实验表明,该方法持续优于先前方法,其中在PTB-XL All分类任务上F1提升超过3%,在SubClass分类任务上提升超过5%。零样本评估下,MERIT在PTB-XL SubClass上AUC和F1分别提升最高达2.66%和2.11%,且在多种分布偏移设置下表现出鲁棒性。此外,利用学习到的心电图表示引导大语言模型生成心电图条件下的临床文本,显著提升文本质量,涵盖ROUGE与METEOR等多个指标。结果表明,MERIT能学习更丰富、更具临床意义的心电图表示,尤其适用于细粒度临床应用。
原文摘要 · Abstract (English)
Electrocardiograms (ECGs) are widely used non-invasive measurements of cardiac activity and play a central role in clinical diagnosis. Recent multimodal approaches align ECG signals with clinical reports to incorporate diagnostic semantics, but clinical reports often fail to preserve the rich physiological structure of ECG waveforms, particularly across multiple levels of abstraction ranging from coarse diagnostic categories to fine-grained morphology. To address this limitation, we formulate ECG representation learning from an information-theoretic perspective and derive a tractable objective that jointly preserves signal structure and integrates clinical semantics. Based on this principle, we propose \textbf{MERIT} (Multimodal ECG Representation via Information Theory), a dual-branch pretraining framework combining masked ECG modeling with ECG--text contrastive alignment. Extensive experiments on PTB-XL and additional benchmarks demonstrate consistent improvements over prior methods, including gains exceeding $3%$ F1 on PTB-XL All and $5%$ F1 on SubClass classification. In zero-shot evaluation, MERIT further improves performance by up to $ +2.66\%$ AUC and $ +2.11\%$ F1 on PTB-XL SubClass, while also demonstrating robustness under multiple distribution-shift settings. Moreover, leveraging the learned ECG representations for ECG-conditioned clinical text generation with large language models improves text quality across several metrics, including ROUGE and METEOR. Together, these results demonstrate that MERIT learns more informative and clinically meaningful ECG representations, particularly for fine-grained clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。