arXiv:2512.12107cs.CV2025-12被引 1

首个基于测量的超声心动图多模态模型,提升心脏图像理解精度。

EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography

  • 构建首个以测量为基础的超声心动图图文数据集,支持多任务训练。
  • 零样本疾病分类AUC达86.5%,视图分类准确率95.1%,性能领先。
  • 适合医疗AI研究者与临床辅助系统开发者使用。

超声心动图是心脏病学中最常用的影像技术,但其解读仍依赖人工且具有多模态特性,需进行视图识别、定量测量、定性评估及指南驱动推理。尽管视觉-语言模型在自然图像和部分医学领域取得成功,但在超声心动图中的应用受限于缺乏大规模、临床相关的图文数据集以及对测量推理的核心缺失。本文提出首个基于测量的多模态超声心动图数据集EchoGround-MIMIC,包含19,065张图像-文本对,来自1,572名患者,涵盖标准化视图、结构化测量、测量引导的描述及指南衍生的疾病标签。在此基础上,我们开发了EchoVLM,引入两种新型预训练目标:(i) 视图感知对比损失,捕捉超声图像的视图依赖结构;(ii) 否定感知对比损失,区分临床上关键的阴性与阳性发现。在5类临床应用共36个任务中(包括多模态疾病分类、图文检索、视图分类、心腔分割和特征点检测),EchoVLM达到当前最优表现,零样本疾病分类AUC为86.5%,视图分类准确率为95.1%。结果表明,临床锚定的多模态预训练可生成可迁移的视觉表征,确立了EchoVLM作为端到端超声心动图解读的基础模型。我们将公开EchoGround-MIMIC数据集及数据清洗代码,促进可复现研究与多模态超声心动图分析的发展。

原文摘要 · Abstract (English)

Echocardiography is the most widely used imaging modality in cardiology, yet its interpretation remains labor-intensive and inherently multimodal, requiring view recognition, quantitative measurements, qualitative assessments, and guideline-based reasoning. While recent vision-language models (VLMs) have achieved broad success in natural images and certain medical domains, their potential in echocardiography has been limited by the lack of large-scale, clinically grounded image-text datasets and the absence of measurement-based reasoning central to echo interpretation. We introduce EchoGround-MIMIC, the first measurement-grounded multimodal echocardiography dataset, comprising 19,065 image-text pairs from 1,572 patients with standardized views, structured measurements, measurement-grounded captions, and guideline-derived disease labels. Building on this resource, we propose EchoVLM, a vision-language model that incorporates two novel pretraining objectives: (i) a view-informed contrastive loss that encodes the view-dependent structure of echocardiographic imaging, and (ii) a negation-aware contrastive loss that distinguishes clinically critical negative from positive findings. Across five types of clinical applications with 36 tasks spanning multimodal disease classification, image-text retrieval, view classification, chamber segmentation, and landmark detection, EchoVLM achieves state-of-the-art performance (86.5% AUC in zero-shot disease classification and 95.1% accuracy in view classification). We demonstrate that clinically grounded multimodal pretraining yields transferable visual representations and establish EchoVLM as a foundation model for end-to-end echocardiography interpretation. We will release EchoGround-MIMIC and the data curation code, enabling reproducibility and further research in multimodal echocardiography interpretation.

超声心动图多模态视觉语言模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。