对比零样本大模型与生理感知CNN在心电图分类中的表现,发现后者更可靠。
Physiology-Aware CNN and Zero-Shot Multimodal LLMs for ECG Image Classification: A Comparative Study
- 用解剖学导联分组设计生理感知CNN,融合多导联特征
- 自研模型内部与外部数据集上AUC达0.92-0.94和0.85-0.86
- 零样本大模型诊断性能接近随机,不适用于临床判断
多模态大语言模型(LLM)被用于解读12导联心电图图像,但其解释缺乏验证。由于心电图理解依赖精确波形形态、导联关系及准确时间间隔测量,与通用图像差异显著。本研究评估了零样本多模态LLM在区分正常与异常心电图上的可靠性,并同时对比基于CNN的模型作为临床基准。将标准12导联心电图生成单页图像,进行二分类任务。测试了三个主流LLM(GPT-5.2、GPT-4.1、Gemini-2.5 Pro),使用固定零样本提示进行多轮测试。同时,开发了一种生理感知的CNN模型,能聚合预定义解剖导联组的特征。该模型与ResNet18、DenseNet121、VGG16基线模型在内部测试集和外部PTB-XL数据集上进行比较。结果表明,基于CNN的模型表现稳定,内部平均ROC-AUC为0.92–0.94,外部为0.85–0.86。所提LeadGroupECG模型在内部性能优于其主干网络,且保持良好的外部泛化能力。相比之下,零样本LLM的判别性能接近随机(ROC-AUC约0.5)。当心电图带有网格校准背景时,精确率-AUC略有提升。尽管多模态大模型可生成合理的心电图描述,但其零样本诊断判别能力仍有限。因此,面向临床、领域特定的架构对心电图人工智能分析仍至关重要。
原文摘要 · Abstract (English)
Multimodal large language models (LLMs) are increasingly adopted to interpret 12-lead ECG images, though the interpretations often lack validation. However, ECG image understanding significantly differs from general images as it depends on precise waveform morphology, lead relationships and accurate interval measurements. This study investigated whether zero-shot multimodal LLMs can reliably distinguish normal and abnormal ECG images and, in parallel, evaluated CNN-based models for clinically grounded references. Standard 12-lead ECG recordings were rendered as single-page images for a binary normal-abnormal classification task. Three prominent LLMs (GPT-5.2, GPT-4.1, and Gemini-2.5 Pro) were tested using a fixed zero-shot prompt across multiple runs. In parallel, a physiology-aware CNN-based model was developed with the capability to aggregate features from the predefined anatomical lead groups. The model was compared with ResNet18, DenseNet121, VGG16 baselines, and all the models were evaluated on an internal test set and external PTB-XL dataset. Across seeds, CNN-based models demonstrated stable discrimination, with average internal ROC-AUC of 0.92-0.94, and external ROC-AUC of 0.85-0.86. The proposed LeadGroupECG model significantly improved over its backbone internally without compromising external generalization. It remained competitive with other baselines, while consistently highlighting anatomical lead-group contributions. In contrast, zero-shot LLM discrimination remained near-chance (ROC-AUC around 0.5). The PR-AUC improved slightly when ECGs used a grid-based calibration background compared with the grid-free ECGs. Although multimodal LLMs can generate reasonable ECG narratives, their zero-shot diagnostic discrimination remains limited. Therefore, clinically framed, domain-specific architectures remain essential for AI-based ECG interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。