arXiv:2602.17531cs.LGcs.AI2026-02

现有心电图表征学习评估需改革,否则研究结果不可靠。

Position: Evaluation of ECG Representations Must Be Fixed

  • 建议扩展评估范围至结构性心脏病和患者预测等临床目标
  • 改进评估方法后,当前最优模型排名发生改变
  • 随机初始化编码器在多数任务上已接近顶尖表现,应作基线

本文指出,当前12导联心电图表征学习的基准测试体系亟需修正,以确保研究进展可靠并契合临床目标。领域内普遍采用的三个公开多标签基准(PTB-XL、CPSC2018、CSN)主要聚焦心律失常与波形形态标签,但心电图实际蕴含更广泛的临床信息。我们主张下游评估应拓展至结构性心脏病与患者层面预测等新目标。此外,针对多标签不平衡场景提出评估最佳实践,应用后发现文献中关于最优表征的结论被改变。更令人意外的是,在多个任务上,随机初始化编码器配合线性评估的表现已达到当前最先进的预训练模型水平。为此,我们基于五个代表性预训练方法,在六个评估场景(三个标准基准、结构性疾病数据集、血流动力学推断、患者预测)中进行实证分析,验证上述观点。

原文摘要 · Abstract (English)

This position paper argues that current benchmarking practice in 12-lead ECG representation learning must be fixed to ensure progress is reliable and aligned with clinically meaningful objectives. The field has largely converged on three public multi-label benchmarks (PTB-XL, CPSC2018, CSN) dominated by arrhythmia and waveform-morphology labels, even though the ECG is known to encode substantially broader clinical information. We argue that downstream evaluation should expand to include an assessment of structural heart disease and patient-level forecasting, in addition to other evolving ECG-related endpoints, as relevant clinical targets. Next, we outline evaluation best practices for multi-label, imbalanced settings, and show that when they are applied, the literature's current conclusion about which representations perform best is altered. Furthermore, we demonstrate the surprising result that a randomly initialized encoder with linear evaluation matches state-of-the-art pre-training on many tasks. This motivates the use of a random encoder as a reasonable baseline model. We substantiate our observations with an empirical evaluation of five representative ECG pre-training approaches across six evaluation settings: the three standard benchmarks, a structural disease dataset, hemodynamic inference, and patient forecasting.

心电图表征学习临床评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。