arXiv:2509.25095eess.SPcs.LG2025-09中稿 · ICLR被引 7

对比8种ECG模型,发现架构比规模更重要,小模型也能超大模型。

Benchmarking ECG FMs: A Reality Check Across Clinical Tasks

  • 在26项临床任务上测试8个ECG基础模型,覆盖12个数据集
  • 小模型ECG-CPC在5类任务中领先,标签效率提升3.3-9倍
  • 不同模型学习到的内部结构差异大,说明路径多样,架构关键

12导联心电图(ECG)是长期使用的诊断工具。然而,心电图机器学习仍呈碎片化,常局限于特定任务或数据集。基础模型(FMs)有望提升泛化能力,但核心问题仍存:哪种架构泛化最好?模型在少量标注下如何扩展?不同模型家族性能差异何在?本研究在包含1,650个回归与分类目标的12个公开数据集上,对8个ECG FMs进行了26项临床任务的基准测试。模型在微调与冻结设置下评估,并分析了不同数据量下的扩展行为。结果显示各领域表现差异显著:在成人ECG解读中,三个FM持续优于强监督基线;而结构化状态空间模型ECG-CPC在7类任务中的5类中领先,表明架构比规模更重要。FM实现3.3至9倍的标签效率提升,但扩展行为因架构而异。表示分析显示,性能相近的模型学习到截然不同的内部结构,暗示存在多条有效的心电图表征路径。总体而言,尽管FM在成人ECG分析中前景可期,但在心脏结构、预后预测和患者特征刻画方面仍存在显著差距。ECG-CPC虽体积小一个数量级却表现优异,挑战了‘大模型=高质量’的假设,凸显架构先验偏置的巨大潜力。

原文摘要 · Abstract (English)

The 12-lead electrocardiogram (ECG) is a long-standing diagnostic tool. Yet machine learning for ECG interpretation remains fragmented, often limited to narrow tasks or datasets. FMs promise broader adaptability, but fundamental questions remain: Which architectures generalize best? How do models scale with limited labels? What explains performance differences across model families? We benchmarked eight ECG FMs on 26 clinically relevant tasks using 12 public datasets comprising 1,650 regression and classification targets. Models were evaluated under fine-tuning and frozen settings, with scaling analyses across dataset sizes. Results show heterogeneous performance across domains: in adult ECG interpretation, three FMs consistently outperformed strong supervised baselines. In contrast, ECG-CPC, a compact structured state-space model, dominated 5 of 7 task categories, demonstrating that architecture matters more than scale. FMs improved label efficiency 3.3-9x over supervised baselines, though scaling behaviors varied across architectures. Representation analysis reveals that models with similar performance learn markedly different internal structures, suggesting multiple viable paths to effective ECG representation. Overall, while FMs show promise for adult ECG analysis, substantial gaps remain in cardiac structure, outcome prediction, and patient characterization. ECG-CPC's strong performance despite being orders of magnitude smaller challenges the assumption that FM quality requires massive scale, highlighting architectural inductive biases as an untapped opportunity.

心电图基础模型架构设计标签效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。