arXiv:2410.19008eess.IVcs.AI2024-10被引 32

用百万级图文数据训练模型,让大模型读懂心电图图像。

Teach Multimodal LLMs to Comprehend Electrocardiographic Images

  • 构建百万样本心电图图文指令数据集,指导大模型理解心电图。
  • 新模型PULSE在四项任务上准确率提升15%至30%,超越通用模型。
  • 适合医疗AI研究者与临床辅助诊断系统开发者参考。

心电图(ECG)是评估心脏状况的重要无创诊断工具。现有自动解读方法泛化能力有限,仅关注少数心脏疾病,且通常依赖原始生理信号,在资源匮乏地区难以使用,而这些地区往往只有打印或数字心电图图像。多模态大语言模型(MLLM)的发展为解决此问题提供了新可能。然而,由于缺乏指令微调数据集和成熟的心电图图像评估基准,其在心电图图像解读中的应用仍面临挑战。为此,我们推出了包含超百万样本的ECGInstruct心电图图像指令微调数据集,覆盖多种数据源的广泛任务。基于此,我们开发了专用于心电图图像理解的PULSE模型。同时,我们构建了涵盖九个数据集、四个关键任务的ECGBench评估基准。实验表明,PULSE达到新基准水平,平均准确率相较通用MLLM提升15%至30%。本工作展示了PULSE在临床实践中提升心电图解读的潜力。

原文摘要 · Abstract (English)

The electrocardiogram (ECG) is an essential non-invasive diagnostic tool for assessing cardiac conditions. Existing automatic interpretation methods suffer from limited generalizability, focusing on a narrow range of cardiac conditions, and typically depend on raw physiological signals, which may not be readily available in resource-limited settings where only printed or digital ECG images are accessible. Recent advancements in multimodal large language models (MLLMs) present promising opportunities for addressing these challenges. However, the application of MLLMs to ECG image interpretation remains challenging due to the lack of instruction tuning datasets and well-established ECG image benchmarks for quantitative evaluation. To address these challenges, we introduce ECGInstruct, a comprehensive ECG image instruction tuning dataset of over one million samples, covering a wide range of ECG-related tasks from diverse data sources. Using ECGInstruct, we develop PULSE, an MLLM tailored for ECG image comprehension. In addition, we curate ECGBench, a new evaluation benchmark covering four key ECG image interpretation tasks across nine different datasets. Our experiments show that PULSE sets a new state-of-the-art, outperforming general MLLMs with an average accuracy improvement of 15% to 30%. This work highlights the potential of PULSE to enhance ECG interpretation in clinical practice.

心电图多模态大模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。