首个基于测量的超声心动图多模态模型,提升心脏图像理解精度。
EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography
- 构建首个以测量为基础的超声心动图图文数据集,支持多任务训练。
- 零样本疾病分类AUC达86.5%,视图分类准确率95.1%,性能领先。
- 适合医疗AI研究者与临床辅助系统开发者使用。
超声心动图是心脏病学中最常用的影像技术,但其解读仍依赖人工且具有多模态特性,需进行视图识别、定量测量、定性评估及指南驱动推理。尽管视觉-语言模型在自然图像和部分医学领域取得成功,但在超声心动图中的应用受限于缺乏大规模、临床相关的图文数据集以及对测量推理的核心缺失。本文提出首个基于测量的多模态超声心动图数据集EchoGround-MIMIC,包含19,065张图像-文本对,来自1,572名患者,涵盖标准化视图、结构化测量、测量引导的描述及指南衍生的疾病标签。在此基础上,我们开发了EchoVLM,引入两种新型预训练目标:(i) 视图感知对比损失,捕捉超声图像的视图依赖结构;(ii) 否定感知对比损失,区分临床上关键的阴性与阳性发现。在5类临床应用共36个任务中(包括多模态疾病分类、图文检索、视图分类、心腔分割和特征点检测),EchoVLM达到当前最优表现,零样本疾病分类AUC为86.5%,视图分类准确率为95.1%。结果表明,临床锚定的多模态预训练可生成可迁移的视觉表征,确立了EchoVLM作为端到端超声心动图解读的基础模型。我们将公开EchoGround-MIMIC数据集及数据清洗代码,促进可复现研究与多模态超声心动图分析的发展。
原文摘要 · Abstract (English)
Echocardiography is the most widely used imaging modality in cardiology, yet its interpretation remains labor-intensive and inherently multimodal, requiring view recognition, quantitative measurements, qualitative assessments, and guideline-based reasoning. While recent vision-language models (VLMs) have achieved broad success in natural images and certain medical domains, their potential in echocardiography has been limited by the lack of large-scale, clinically grounded image-text datasets and the absence of measurement-based reasoning central to echo interpretation. We introduce EchoGround-MIMIC, the first measurement-grounded multimodal echocardiography dataset, comprising 19,065 image-text pairs from 1,572 patients with standardized views, structured measurements, measurement-grounded captions, and guideline-derived disease labels. Building on this resource, we propose EchoVLM, a vision-language model that incorporates two novel pretraining objectives: (i) a view-informed contrastive loss that encodes the view-dependent structure of echocardiographic imaging, and (ii) a negation-aware contrastive loss that distinguishes clinically critical negative from positive findings. Across five types of clinical applications with 36 tasks spanning multimodal disease classification, image-text retrieval, view classification, chamber segmentation, and landmark detection, EchoVLM achieves state-of-the-art performance (86.5% AUC in zero-shot disease classification and 95.1% accuracy in view classification). We demonstrate that clinically grounded multimodal pretraining yields transferable visual representations and establish EchoVLM as a foundation model for end-to-end echocardiography interpretation. We will release EchoGround-MIMIC and the data curation code, enabling reproducibility and further research in multimodal echocardiography interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。