arXiv:2607.03009cs.LG2026-07

预训练模型对罕见心脏病检测帮助有限,主要提升训练稳定性而非临床泛化能力。

Do ECG Foundation Models Transfer to Rare Cardiac Diseases? Evidence from Brugada Syndrome Detection

  • 用多种预训练模型在两个数据集上测试,比较从零训练、线性探测和微调效果
  • 在3%数据下仅轻微提升性能(AUC增0.055),跨中心零样本迁移表现接近随机
  • 高容量模型依赖预训练才能收敛,但小型模型无需预训练也能达到最优

背景:大规模无标签生理数据训练的基底模型(FMs)为医学人工智能提供了新范式,但其对罕见心电图表型的可转移临床表征能力尚未验证。本研究系统评估了九种公开可用的ECG基底模型在Brugada综合征检测中的表现,使用BrSwiss队列(294例患者,87例病例)和独立外部HUCA队列(363例患者,76例病例),采用三种策略(从零训练、线性探测、全量微调),涵盖3%数据删减和零样本跨中心迁移。结果:对于无法从零收敛的高容量架构,预训练带来显著性能提升(最高AUC增0.411,p<0.05),但小型架构在仅有标注数据时已可收敛且无需预训练。在完整BrSwiss数据集上,最佳微调模型ECG-CPC(AUC=0.962)仅略优于从零训练的监督基线(AUC=0.932;p=0.091)。在匹配训练集大小下,BrSwiss-3%的数据效率优势(AUC增0.055,p<0.01)未在HUCA队列复现。零样本跨中心迁移中,基于基底模型的流程并未优于监督基线,全部接近随机水平。结论:对于Brugada综合征检测,基底模型预训练更多是优化稳定性而非可转移的临床知识,挑战了大规模预训练必然蕴含临床意义表征的假设,强调模型架构与数据域对齐的关键作用。

原文摘要 · Abstract (English)

Background: Foundation models (FMs) trained on large-scale unlabeled physiological data have emerged as a promising paradigm for medical artificial intelligence. Their ability to capture clinically meaningful, transferable representations for rare diseases remains largely unproven. This study investigates whether FM pre-training provides genuine clinical generalization benefits beyond improved optimization for rare electrocardiographic (ECG) phenotypes. Methods: We systematically evaluated nine publicly available ECG FMs for Brugada syndrome detection on the BrSwiss cohort (294 patients, 87 cases) and the independent external HUCA cohort (363 patients, 76 cases), under three strategies (from-scratch training, linear probing, full fine-tuning) across several configurations, including a 3% data ablation and zero-shot cross-site transfers. Results: Pre-training was necessary for high-capacity architectures unable to converge from scratch (AUC gain up to 0.411, p < 0.05), but gave no significant gain for compact architectures already converged on labeled data alone. On full BrSwiss, the best fine-tuned FM (ECG-CPC, AUC = 0.962) only marginally exceeded the strongest supervised baseline (ECG-CPC from scratch, AUC = 0.932; p = 0.091). At matched training-set size, the data-efficiency advantage on BrSwiss-3% (AUC gain = 0.055, p < 0.01) did not replicate on HUCA. Under zero-shot cross-site transfer, FM-based pipelines did not generalize better than supervised baselines, all approaching chance-level performance. Conclusion: For Brugada syndrome detection, FM pre-training is mechanical rather than semantic, providing optimization stability rather than transferable clinical knowledge. These findings challenge the assumption that large-scale pre-training inherently encodes clinically meaningful representations, highlighting the central role of model architecture and data-domain alignment.

ECG分析基底模型罕见病检测跨中心迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。