arXiv:2607.25117eess.IVcs.AI2026-07

研究心电图图像分类中模型是否依赖真实波形或人为干扰信号。

Analysis of the Shortcut Learning and Clever Hans Effect in CNN based ECG Image Classification

  • 构建六种可控图像特征集测试模型对波形信息的依赖程度。
  • 发现模型在移除波形后仍保持高准确率,存在明显捷径学习现象。
  • 适合关注AI医疗可解释性与临床可信度的研究者阅读。

基于卷积神经网络的心电图图像分类模型可能通过利用非生理学视觉线索而非心电图波形形态实现高精度。由于深度学习模型具有黑箱特性,其高预测性能往往难以转化为临床信任、可解释性及可操作决策。本研究在公开的心电图图像数据集上,分析了捷径学习与聪明汉斯效应。我们创建了六种图像衍生特征集:FS1为原始完整心电图图像,FS2为仅包含波形的裁剪图像,FS3为波形遮蔽的元数据图像,FS4为心肌梗死类别的红色箭头伪影图像,FS5为异常心律类别的对比增强图像,FS6为正常类别的高斯模糊图像。这些受控表示用于测试当波形信息被移除或引入特定类别人工伪影时,分类性能是否依然保持。通过计算捷径保留分数、预测一致性及置信度差异,评估模型学习模式的透明度。同时展示平均集成梯度与遮挡敏感性测试结果,以检验模型注意力是否集中在心电图相关波形区域,还是非临床伪影上。性能变化与归因模式用于识别潜在的聪明汉斯行为。本研究评估心电图图像分类器是否学习到临床有意义的波形特征,还是依赖报告布局、元数据、对比度、模糊或人工标记等捷径线索。

原文摘要 · Abstract (English)

Deep learning models for ECG image classification may achieve high accuracy by exploiting non-physiological visual cues instead of ECG waveform morphology. Given the black-box nature of deep learning models, their promise of high predictive performance often remains insufficiently translated into clinical or real-world trust, interpretability, and actionable decision-making. In this study, we examine shortcut learning and Clever Hans effect in a publicly available ECG image dataset using convolutional neural networks. In process we have created six image-derived feature sets (FSs), FS1: raw full ECG images, FS2: cropped waveform-only images, FS3: waveform-masked metadata images, FS4: red-arrow artifact images for the myocardial infarction class, FS5: contrast-enhanced images for the abnormal heartbeat class and FS6: Gaussian-blurred images for the normal class. These controlled representations were used to test whether classification performance persists when waveform information is removed or when artificial class-specific artifacts are introduced. Shortcut retention score, prediction consistency and confidence divergence across Feature-Set Representations have been calculated to assess the transparency about the learning pattern. Along with factual results, average Integrated Gradients and occlusion sensitivity test results are presented to inspect whether model attribution focused on ECG-relevant waveform regions or on non-clinical artifacts. Performance changes across feature sets and attribution patterns were used to identify potential Clever Hans behavior. This study evaluates whether ECG image classifiers learn clinically meaningful morphology or shortcut cues introduced by report layout, metadata, contrast, blur, or artificial markers.

心电图可解释性捷径学习CNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。