用临床指南评估心电图模型解释,发现多数方法只看信号强弱,忽略真正关键的低幅区域。
Beyond Local Inspection: Global, Guideline-Grounded Evaluation of Post-hoc XAI Methods for ECG Classification

- 基于临床指南构建全局评估框架,对比解释与诊断相关区域的匹配度。
- 13种梯度方法中9种在至少一种情况下表现低于随机水平,平均斯皮尔曼相关达0.69。
- 适合医疗AI可解释性研究者,尤其关注心脏疾病模型的可靠性验证。
可解释人工智能(XAI)用于检验模型是否依赖有意义的模式,但看似合理的单个预测解释可能系统性地误导模型行为判断。这在医学领域尤为严重,因为模型可能依赖无关信号特征而非疾病特异性模式而未被察觉。本文以心电图(ECG)数据为例,利用临床指南明确诊断相关信号区域,提出一种全局、基于指南的评估框架,通过聚合心跳层面的解释来评估其与临床定义兴趣区域的一致性。在PTB-XL数据集上训练的四个二分类器,评估了13种基于梯度的方法在两类模式下的表现:低幅段和高幅QRS形态。结果表明,从计算机视觉迁移来的方法普遍存在系统性失败:其解释常跟随信号幅度而非临床相关性,平均斯皮尔曼相关高达0.69,导致忽视诊断决定性的低幅区域。对于缺血性病变,LRP-ε仅将4.6%的重要性分配给ST段,而LRP-SIGN为63.8%。13种方法中有9种在至少一种条件下表现低于随机水平,表明其在不同模式下可靠性不一致。这些发现说明,基于领域知识的全局评估能揭示样本级热图无法察觉的系统性解释缺陷。
原文摘要 · Abstract (English)
Explainable AI (XAI) is used to assess whether artificial intelligence models rely on meaningful patterns, yet explanations that appear plausible for individual predictions may systematically misrepresent model behavior. This is particularly problematic in medicine, where models may rely on irrelevant signal characteristics rather than disease-specific patterns without being recognizable. We address this challenge using electrocardiogram (ECG) data, for which clinical guidelines provide explicit knowledge about diagnostically relevant signal regions. We introduce a global, guideline-grounded framework that aggregates explanations across heartbeats to evaluate them against clinically defined regions of interest. Using four binary classifiers trained on PTB-XL, we assess 13 gradient-based methods across two categories of patterns: low-amplitude segments and high-amplitude QRS morphology. Our results reveal a systematic failure of methods transferred from computer vision. Their explanations often follow signal amplitude rather than clinical relevance, with mean Spearman correlations up to 0.69, leading them to overlook diagnostically decisive low-amplitude regions. For ischemia, LRP-$ε$ assigns only 4.6% of relevance to the ST segment, compared with 63.8% for LRP-SIGN. Nine of 13 methods fall below chance for at least one condition, indicating inconsistent reliability across patterns. These findings show that global, domain-grounded evaluation can uncover systematic explanation failures not obvious from sample-level heatmaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。