用标准化自编码器评估心电图大模型的可解释性,揭示不同模型表现差异。
ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

- 采用匹配容量的稀疏自编码器作为统一测量工具
- 发现六款模型在可解释性上各有优劣,重建精度与临床可读性不一致
- 适合关注模型内部机制、临床可解释性的研究人员使用
现有心电图大模型评测主要关注下游预测性能,难以反映其内部表征是否可被准确分解、临床解读或跨分析复现。我们提出ECG-InterpBench,一个系统评估心电图大模型表征可解释性的基准。该基准使用稀疏自编码器作为标准化测量工具,并匹配各模型间的容量以实现可控比较。我们评估了六款冻结的心电图大模型,覆盖五种标准编码器深度、五种匹配字典宽度和三个随机种子,构建了包含450个单元的可解释性图谱,其中75个为六模型精确匹配的对比块。基准评估了稀疏重建保真度、单特征可访问性、49项临床相关心电图指标的覆盖度以及跨种子特征重现性等维度。进一步量化了患者采样不确定性、深度与种子依赖性变异及稀疏参数敏感性。结果显示,不同心电图大模型具有显著不同的可解释性特征。在MIMIC-IV-ECG上的复现验证表明,重建保真度与临床可访问性识别出不同领先模型。该基准附带可执行代码、标准化清单、单元级指标与可复现性审计。ECG-InterpBench通过容量控制与可复现框架,补足了以性能为中心的评测,支持从多维度对比心电图大模型的表征可解释性。
原文摘要 · Abstract (English)
Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically interpreted, or reproduced across independent analyses. We introduce ECG-InterpBench, a benchmark designed to systematically evaluate the interpretability of ECG foundation-model representations. ECG-InterpBench uses sparse autoencoders as standardized measurement instruments and matches their capacity across models to enable controlled comparisons. We evaluate six frozen ECG foundation models across five standardized encoder depths, five matched dictionary widths, and three random seeds, producing a 450-cell interpretability atlas comprising 75 exactly matched six-model comparison blocks. The benchmark evaluates complementary dimensions of representation interpretability, including sparse reconstruction fidelity, single-feature accessibility and coverage of 49 clinically meaningful ECG measurements, and cross-seed feature reproducibility. The evaluation further quantifies patient-sampling uncertainty, depth- and seed-dependent variation, and sensitivity to the sparsity parameterization. The benchmark reveals that ECG foundation models exhibit distinct interpretability profiles. A matched replication on MIMIC-IV-ECG confirms that reconstruction fidelity and clinical accessibility identify different leading models. The benchmark is accompanied by executable evaluation code, standardized manifests, cell-level metrics, and reproducibility audits. ECG-InterpBench complements performance-centered ECG benchmarks by providing a capacity-controlled and reproducible framework for comparing ECG foundation models across distinct dimensions of representation interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。