对比了心电图多模态融合的两种方式,发现中间融合更准且可解释。
Explainable Deep Neural Network for Multimodal ECG Signals: Intermediate vs Late Fusion
- 用时间、频率、时频域信号做多模态融合,比较中间与晚期融合效果
- 中间融合最高准确率达97%,比单独模型提升显著(Cohen's d > 0.8)
- 通过显著性图揭示模型决策依据,适合医疗AI可信应用
单模态深度学习模型存在过拟合和泛化能力弱的问题,促使研究者重新关注多模态融合策略。多模态深度神经网络(MDNN)能整合不同数据域,有望实现更稳健、准确的预测。然而,在心电图(ECG)相关心血管疾病(CVD)分类等高风险临床场景中,特征级融合(中间融合)与决策级融合(晚期融合)的最优策略仍不明确。本研究在时间、频率、时频三个域的ECG信号上对比分析两种融合策略。实验表明,中间融合持续优于晚期融合,最高准确率达到97%,相较独立模型的效应量Cohen's d > 0.8,相较晚期融合的效应量d = 0.40。通过显著性图进行可解释性分析,发现两类模型均与离散化的心电图信号一致。利用互信息(MI)验证了各分类下离散化信号与显著性图之间的统计依赖关系。所提出的基于心电图域的多模态模型具备更强的预测性能与可解释性,优于现有先进方法,对医疗AI应用具有重要意义。
原文摘要 · Abstract (English)
The limitations of unimodal deep learning models, particularly their tendency to overfit and limited generalizability, have renewed interest in multimodal fusion strategies. Multimodal deep neural networks (MDNN) have the capability of integrating diverse data domains and offer a promising solution for robust and accurate predictions. However, the optimal fusion strategy, intermediate fusion (feature-level) versus late fusion (decision-level) remains insufficiently examined, especially in high-stakes clinical contexts such as ECG-based cardiovascular disease (CVD) classification. This study investigates the comparative effectiveness of intermediate and late fusion strategies using ECG signals across three domains: time, frequency, and time-frequency. A series of experiments were conducted to identify the highest-performing fusion architecture. Results demonstrate that intermediate fusion consistently outperformed late fusion, achieving a peak accuracy of 97 percent, with Cohen's d > 0.8 relative to standalone models and d = 0.40 compared to late fusion. Interpretability analyses using saliency maps reveal that both models align with the discretized ECG signals. Statistical dependency between the discretized ECG signals and corresponding saliency maps for each class was confirmed using Mutual Information (MI). The proposed ECG domain-based multimodal model offers superior predictive capability and enhanced explainability, crucial attributes in medical AI applications, surpassing state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。