构建超大规模心电图数据集,助力AI精准诊断
CODE-II: A large-scale dataset for artificial intelligence in ECG analysis
- 基于巴西远程医疗网络收集270万份心电图,标准化标注
- 66类临床诊断标签,公开1.5万患者子集供研究使用
- 预训练模型在外部数据集上表现优于更大规模数据训练模型
基于数据驱动的心电图分析方法快速发展,但标注质量、数据规模与覆盖范围仍存局限。本文提出CODE-II,一个来自巴西米纳斯吉拉斯州远程医疗网络(TNMG)的大型真实世界心电图数据集,包含2,735,269份12导联心电图,来自2,093,807名成年患者。每份检查均采用标准化诊断标准并经心脏病专家审核。CODE-II具有66个由心脏病专家参与定义的临床有意义诊断类别,广泛用于远程医疗实践。此外,我们提供公开子集CODE-II-open(15,000名患者)及非重叠测试集CODE-II-test(8,475份检查,经多位心脏病专家盲评)。在CODE-II上预训练的神经网络在外部基准(PTB-XL和CPSC 2018)上展现出优异的迁移性能,优于在更大规模数据集上训练的模型。
原文摘要 · Abstract (English)
Data-driven methods for electrocardiogram (ECG) interpretation are rapidly progressing. Large datasets have enabled advances in artificial intelligence (AI) based ECG analysis, yet limitations in annotation quality, size, and scope remain major challenges. Here we present CODE-II, a large-scale real-world dataset of 2,735,269 12-lead ECGs from 2,093,807 adult patients collected by the Telehealth Network of Minas Gerais (TNMG), Brazil. Each exam was annotated using standardized diagnostic criteria and reviewed by cardiologists. A defining feature of CODE-II is a set of 66 clinically meaningful diagnostic classes, developed with cardiologist input and routinely used in telehealth practice. We additionally provide an open available subset: CODE-II-open, a public subset of 15,000 patients, and the CODE-II-test, a non-overlapping set of 8,475 exams reviewed by multiple cardiologists for blinded evaluation. A neural network pre-trained on CODE-II achieved superior transfer performance on external benchmarks (PTB-XL and CPSC 2018) and outperformed alternatives trained on larger datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。