120万份心电图数据构建基准,验证公开数据训练心电图大模型的有效性。
OpenECG: Benchmarking ECG Foundation Models with Public 1.2 Million Records
- 用自监督学习在9个中心的120万份心电图上训练模型
- BYOL和MAE在泛化性能上优于SimCLR,且数据量达60%-70%时性能饱和
- 证明公开数据可媲美私有数据,适合医疗AI研究者参考
本研究提出OpenECG,一个包含来自九个中心的120万份12导联心电图记录的大规模基准,用于评估基于公开数据集训练的心电图基础模型(ECG-FMs)。我们采用ResNet-50与Vision Transformer架构,测试了三种自监督学习方法(SimCLR、BYOL、MAE),通过留一数据集外实验和数据缩放分析评估模型泛化能力。结果表明,在多样化数据上预训练显著提升泛化性能,其中BYOL和MAE优于SimCLR,凸显特征一致性与生成式学习相较于对比学习的优势。数据缩放实验显示,BYOL和MAE在总数据量达到60%-70%时性能趋于饱和,而SimCLR需要更多数据。这些发现表明,公开心电图数据足以训练出鲁棒的ECG-FMs,其性能可媲美甚至超越专有数据集,为可扩展、临床有意义的心电图人工智能分析铺平道路。
原文摘要 · Abstract (English)
This study introduces OpenECG, a large-scale benchmark of 1.2 million 12-lead ECG recordings from nine centers, to evaluate ECG foundation models (ECG-FMs) trained on public datasets. We investigate three self-supervised learning methods (SimCLR, BYOL, MAE) with ResNet-50 and Vision Transformer architectures, assessing model generalization through leave-one-dataset-out experiments and data scaling analysis. Results show that pre-training on diverse datasets significantly improves generalization, with BYOL and MAE outperforming SimCLR, highlighting the efficacy of feature-consistency and generative learning over contrastive approaches. Data scaling experiments reveal that performance saturates at 60-70% of total data for BYOL and MAE, while SimCLR requires more data. These findings demonstrate that publicly available ECG data can match or surpass proprietary datasets in training robust ECG-FMs, paving the way for scalable, clinically meaningful AI-driven ECG analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。