arXiv:2509.10369cs.LGcs.AI2025-09

对比学习模型性能受训练数据分布影响,多样性反而降低泛化能力。

Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms

  • 用多中心、跨人群数据预训练心电图模型,发现数据分布决定下游表现
  • 跨区域数据提升本地准确率,但导致跨区域泛化能力下降
  • 提出新策略保留组内一致性,增强模型对未知数据的鲁棒性

对比学习是广泛采用的自监督预训练方法,但其对队列构成的依赖尚未充分探索。我们构建了基于患者增强心电图的对比学习基础模型(CAPE),在来自三大洲(北美洲、南美洲、亚洲)的四个队列(n = 5,203,352)上进行预训练,并系统评估队列人口统计学特征、健康状况及群体多样性对下游预测任务的影响,同时包含来自欧洲的两个额外队列。结果表明,下游性能取决于预训练队列的分布特性,包括人口和健康状态。尽管多中心、多样化队列可提高分布内准确率,却会因编码队列特异性伪影而削弱分布外(OOD)泛化能力。为此,我们提出在分布内批次(IDB)策略,在预训练中保持组内一致性,显著提升分布外鲁棒性。本研究为构建临床公平且具备泛化能力的基础模型提供了关键洞见。

原文摘要 · Abstract (English)

Contrastive learning is a widely adopted self-supervised pretraining strategy, yet its dependence on cohort composition remains underexplored. We present Contrasting by Patient Augmented Electrocardiograms (CAPE) foundation model and pretrain on four cohorts (n = 5,203,352), from diverse populations across three continents (North America, South America, Asia). We systematically assess how cohort demographics, health status, and population diversity influence the downstream performance for prediction tasks also including two additional cohorts from another continent (Europe). We find that downstream performance depends on the distributional properties of the pretraining cohort, including demographics and health status. Moreover, while pretraining with a multi-centre, demographically diverse cohort improves in-distribution accuracy, it reduces out-of-distribution (OOD) generalisation of our contrastive approach by encoding cohort-specific artifacts. To address this, we propose the In-Distribution Batch (IDB) strategy, which preserves intra-cohort consistency during pretraining and enhances OOD robustness. This work provides important insights for developing clinically fair and generalisable foundation models.

心电图对比学习泛化能力基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。