首个同步信号、图像、文本的多模态心电图数据集,助力可解释AI诊断
MEETI: A Multimodal ECG Dataset from MIMIC-IV-ECG with Signals, Images, Features and Interpretations
- 整合原始波形、图像、参数与大模型生成的文本解读,四维对齐
- 包含每搏级定量参数,支持精细分析与模型可解释性
- 适合心血管AI研究者构建可解释的多模态诊断系统
心电图(ECG)在现代心血管诊疗中起基础作用,可无创诊断心律失常、心肌缺血和传导障碍。尽管机器学习已在心电图解析中达到专家水平,但临床可用的多模态AI系统发展受限,主要因缺乏同时包含原始信号、诊断图像和解释文本的公开数据集。现有大多数心电图数据集仅提供单模态或双模态数据,难以构建能融合多种心电信息的现实系统。为此,我们推出MEETI(MIMIC-IV-Ext ECG-Text-Image),首个大规模心电图数据集,同步包含原始波形、高分辨率绘图图像及大语言模型生成的详细文本解读。此外,MEETI还提取了每导联的搏动级定量参数,支持细粒度分析与模型可解释性。每个记录在四个组件间严格对齐:(1)原始心电波形,(2)对应绘图图像,(3)提取特征参数,(4)详细解释文本,通过唯一标识符实现一致对齐。该统一结构支持基于Transformer的多模态学习,促进对心脏健康的细粒度、可解释推理。通过连接传统信号分析、图像解读与语言理解,MEETI为下一代可解释的多模态心血管AI奠定了坚实基础,为研究社区提供了全面的基准,用于开发与评估基于心电图的AI系统。
原文摘要 · Abstract (English)
Electrocardiogram (ECG) plays a foundational role in modern cardiovascular care, enabling non-invasive diagnosis of arrhythmias, myocardial ischemia, and conduction disorders. While machine learning has achieved expert-level performance in ECG interpretation, the development of clinically deployable multimodal AI systems remains constrained, primarily due to the lack of publicly available datasets that simultaneously incorporate raw signals, diagnostic images, and interpretation text. Most existing ECG datasets provide only single-modality data or, at most, dual modalities, making it difficult to build models that can understand and integrate diverse ECG information in real-world settings. To address this gap, we introduce MEETI (MIMIC-IV-Ext ECG-Text-Image), the first large-scale ECG dataset that synchronizes raw waveform data, high-resolution plotted images, and detailed textual interpretations generated by large language models. In addition, MEETI includes beat-level quantitative ECG parameters extracted from each lead, offering structured parameters that support fine-grained analysis and model interpretability. Each MEETI record is aligned across four components: (1) the raw ECG waveform, (2) the corresponding plotted image, (3) extracted feature parameters, and (4) detailed interpretation text. This alignment is achieved using consistent, unique identifiers. This unified structure supports transformer-based multimodal learning and supports fine-grained, interpretable reasoning about cardiac health. By bridging the gap between traditional signal analysis, image-based interpretation, and language-driven understanding, MEETI established a robust foundation for the next generation of explainable, multimodal cardiovascular AI. It offers the research community a comprehensive benchmark for developing and evaluating ECG-based AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。