arXiv:2607.20814cs.AI2026-07中稿 · CITA 2026

用临床指南增强心脏诊断报告,让AI解释更可信。

Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs

论文配图:Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs
图 1 · 摘自论文原文
  • 将权威心电图指南转化为结构化知识块,嵌入生成流程
  • 报告与标准术语对齐,BERTScore提升至0.953
  • 适合需要可解释性的心脏病学AI应用

心电图(ECG)是心脏评估的核心工具,但深度学习模型在临床应用中受限于可解释性差和大语言模型(LLM)的幻觉风险。现有CNN+Grad-CAM+多模态LLM框架虽能生成报告,但解释往往未能严格遵循既定诊断标准,降低可信度与可复现性。本文提出一种引导式多模态框架,将精选的临床知识显式锚定于报告生成过程。首先,卷积神经网络(CNN)与Grad-CAM从12导联心电图图像中生成类别概率与类特定热力图;同时,权威心电图教材与指南材料被离线提炼为结构化《心电图解读指南》,作为每例样本固定的外部知识块。在输入包括心电图图像、Grad-CAM叠加图、CNN生成的事实包及注入的指南基础上,多模态LLM生成符合指南术语与判断标准的结构化诊断报告。在完整PTB-XL测试集上的实验表明,引入引导知识显著提升了报告语义质量与感知一致性,同时保持竞争性分类性能。特别地,生成摘要的平均BERTScore从基准的0.818提升至0.953,表明与参考报告的语义对齐度大幅提升。结果表明,在多模态提示流程中注入提炼后的解读指南,是减少幻觉、提升基于LLM的心电图解释临床合理性的可行路径,推动可解释心脏诊断向真实世界部署迈进。

原文摘要 · Abstract (English)

The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning models remains con- strained by limited interpretability and the hallucination risk of large language models (LLMs). Existing CNN+Grad-CAM+multimodal LLM frameworks can generate ECG reports, but their explanations are often only weakly grounded in established diagnostic criteria, reducing trust- worthiness and reproducibility. We propose a guide-grounded multimodal framework that explicitly anchors report generation in curated clinical knowledge. A convolutional neural network (CNN) and Grad-CAM first produce class probabilities and class-specific heatmaps from 12-lead ECG images. In parallel, authoritative ECG textbooks and guideline materials are distilled offline into a structured ECG Interpretation Guide, which is injected as a fixed knowledge block for every sample. Conditioned on the ECG image, Grad-CAM overlay, CNN-derived fact pack, and the in- jected guide, a multimodal LLM generates structured diagnostic reports with guideline-consistent terminology and criteria usage. Experiments on the full PTB-XL test set demonstrate that guide grounding improves se- mantic quality and perceived consistency of generated reports while pre- serving competitive classification performance. In particular, our method increases the average BERTScore of generated impressions from 0.818 to 0.953 relative to a strong CNN+Grad-CAM+MLLM baseline, indicat- ing closer alignment with reference reports. These findings suggest that injecting a distilled interpretation guide into the multimodal prompting pipeline offers a practical pathway to reduce hallucinations and enhance the clinical plausibility of LLM-based ECG explanations, bringing ex- plainable cardiac diagnosis closer to real-world deployment.

心电图可解释AI多模态临床指南

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。