arXiv:2607.25244cs.AI2026-07

将心电图大模型拆解为可解释的生理概念词典,提升模型透明度。

CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models

论文配图:CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models
图 1 · 摘自论文原文
  • 用稀疏自编码器将900万心电信号分解为8192个生理原子。
  • 原子在识别心律失常等临床特征时准确率达0.88(AUROC),优于原始嵌入。
  • 可定位每个预测的生理依据,适合医学AI可解释性研究者使用。

心电图基础模型在临床任务中迁移效果好,但其表征中蕴含的生理知识难以理解。我们提出CADENCE,将心电图基础模型分解为可人类解读、可查询的生理概念词典。通过批处理TopK稀疏自编码器,将超过九百万心电信号的第6层嵌入分解为8192个稀疏心脏原子。这些原子与临床表型和波形形态的对应关系优于单个密集嵌入维度,能够恢复心律失常、传导异常、梗死及复极化模式、心腔与轴向发现,以及导联和心跳相位特异的波形基元。在第6层,最优原子对临床表型和形态的平均AUROC分别达到0.88和0.90,高于最优密集维度的0.78和0.83。稀疏原子探针在表型、形态和年龄预测上表现匹配或超越密集探针,且能将每个预测归因于少数可解释原子;表型预测的AUROC从0.93提升至0.95。原子空间几何结构再现了生理上合理的关联,针对性原子消融可选择性改变冻结下游输出。自动化LLM流水线生成并定量验证原子描述,通过预测保留激活实现。在独立外部心电数据集上,CADENCE能恢复重叠概念并保持一致的表型预测性能。CADENCE提供了一个可扩展的框架,用于发现和审计心电图基础模型中的生理知识。

原文摘要 · Abstract (English)

Foundation models for 12-lead electrocardiograms (ECGs) transfer well across clinical tasks, but the physiological knowledge encoded in their representations remains opaque. We present CADENCE, a framework that decomposes an ECG foundation model into a human-interpretable, queryable dictionary of physiological concepts. Using a BatchTopK sparse autoencoder, CADENCE factorizes Layer-6 embeddings from more than nine million ECG tokens into 8,192 sparse cardiac atoms. These atoms align better than individual dense embedding dimensions with clinical phenotypes and waveform morphology, recovering arrhythmias, conduction abnormalities, infarction and repolarization patterns, chamber and axis findings, and lead- and beat-phase-specific waveform primitives. At Layer 6, the best atoms achieve mean AUROCs of 0.88 for clinical phenotypes and 0.90 for morphology, versus 0.78 and 0.83 for the best dense dimensions. Sparse atom probes match or outperform dense probes for phenotype, morphology, and age prediction while attributing each prediction to a small set of interpretable atoms; phenotype AUROC improves from 0.93 to 0.95. Atom-space geometry recovers physiologically coherent relationships, and targeted atom ablation selectively changes frozen downstream outputs. An automated LLM pipeline generates and quantitatively validates atom descriptions by predicting held-out activations. On independent external ECG datasets, CADENCE recovers overlapping concepts and maintains consistent phenotype-prediction performance. CADENCE provides a scalable framework for discovering and auditing the physiological knowledge encoded by ECG foundation models.

心电图可解释性基础模型稀疏编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。