对比多种自监督方法,发现结构化状态空间模型在心电图建模中表现最优。
Pretraining Strategies and Scaling for ECG Foundation Models: A Systematic Study

- 采用对比与非对比自监督学习,系统评估五种预训练策略
- 数据量达1100万样本时性能仍持续提升,证明规模效应显著
- 结构化状态空间模型优于Transformer和CNN,其先验知识是关键
专用基础模型正逐步在多个医学子领域涌现,但预训练方法及参数规模随预训练数据集增大时的表现仍缺乏系统性、可比性的评估。本文聚焦于全球最广泛采集的生理时间序列之一——心电图(ECG)数据,全面评估了五种不同的对比与非对比自监督学习目标在ECG基础模型中的应用,并研究其在最大1100万样本(仅来自公开数据源)下的扩展行为。预训练策略对下游任务表现具有显著且一致的影响,其中对比预测编码(略优于JEPA)在跨多种临床任务中生成最具迁移能力的表征。多数预训练目标在1100万样本规模下仍表现出性能持续提升。同时,我们在所有预训练方法中比较不同模型架构,发现结构化状态空间模型明显优于Transformer和卷积神经网络。我们推测,结构化状态空间模型的强先验偏置,而非单纯依赖预训练规模,是实现高效心电图表示学习的主要驱动力,这对该领域及其他生理信号领域的未来基础模型开发具有重要意义。
原文摘要 · Abstract (English)
Specialized foundation models are beginning to emerge in various medical subdomains, but pretraining methodologies and parametric scaling with the size of the pretraining dataset are rarely assessed systematically and in a like-for-like manner. This work focuses on foundation models for electrocardiography (ECG) data, one of the most widely captured physiological time series world-wide. We present a comprehensive assessment of pretraining methodologies, covering five different contrastive and non-contrastive self-supervised learning objectives for ECG foundation models, and investigate their scaling behavior with pretraining dataset sizes up to 11M input samples, exclusively from publicly available sources. Pretraining strategy has a meaningful and consistent impact on downstream performance, with contrastive predictive coding (slightly ahead of JEPA) yielding the most transferable representations across diverse clinical tasks. Scaling pretraining data continues to yield meaningful improvements up to 11M samples for most objectives. We also compare model architectures across all pretraining methodologies and find evidence for a clear superiority of structured state space models compared to transformers and CNN models. We hypothesize that the strong inductive biases of structured state space models, rather than pretraining scale alone, are the primary driver of effective ECG representation learning, with important implications for future foundation model development in this and potentially other physiological signal domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。