arXiv:2602.10553cs.LGcs.AI2026-02被引 1

用改进的sigmoid损失提升多标签心电图分类效果

Contrastive Learning for Multi Label ECG Classification with Jaccard Score Based Sigmoid Loss

  • 基于CLIP设计多标签心电图编码器,引入杰卡德分数优化损失函数
  • 在真实医院数据上实现更高准确率,关键发现包括嵌入维度增加与随机裁剪有效缓解数据偏移
  • 适合想做医疗多模态模型、心电图分析的开发者参考

近年来,大语言模型推动了多模态医学AI的发展。尽管MedGemini等模型在USMLE MM等视觉问答任务中表现优异,但在心电图(ECG)任务上的表现仍有限,部分模型如MedGemma甚至不支持ECG数据。解读心电图本身具有挑战性,诊断准确性受医生经验影响。虽然超声心动图提供丰富信息,但依赖专业设备和人员,普及性受限。本研究聚焦于利用真实医院数据构建稳健的心电图编码器用于多模态预训练。采用基于CLIP的SigLIP模型,结合基于sigmoid的损失函数实现多标签预测,并提出针对心电图多标签特性的改进损失函数。实验表明,融合医学知识的语言模型与改进损失显著提升多标签心电图分类性能。为进一步优化,我们增加嵌入维度并应用随机裁剪以缓解数据漂移。逐标签分析揭示了不同心电图特征的预测难易程度。本研究为开发利用心电图数据的医疗模型提供了基础框架。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have enabled the development of multimodal medical AI. While models such as MedGemini achieve high accuracy on VQA tasks like USMLE MM, their performance on ECG based tasks remains limited, and some models, such as MedGemma, do not support ECG data at all. Interpreting ECGs is inherently challenging, and diagnostic accuracy can vary depending on the interpreter's experience. Although echocardiography provides rich diagnostic information, it requires specialized equipment and personnel, limiting its availability. In this study, we focus on constructing a robust ECG encoder for multimodal pretraining using real world hospital data. We employ SigLIP, a CLIP based model with a sigmoid based loss function enabling multi label prediction, and introduce a modified loss function tailored to the multi label nature of ECG data. Experiments demonstrate that incorporating medical knowledge in the language model and applying the modified loss significantly improve multi label ECG classification. To further enhance performance, we increase the embedding dimensionality and apply random cropping to mitigate data drift. Finally, per label analysis reveals which ECG findings are easier or harder to predict. Our study provides a foundational framework for developing medical models that utilize ECG data.

心电图分析多标签学习医疗多模态对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。