arXiv:2605.24159cs.CV2026-05

构建首个大规模超声心动图问答数据集,助力新手操作时实时获取诊断指导。

EchoVQA: Enabling Conversational Assistance for Point-of-Care Cardiac Ultrasound

论文配图:EchoVQA: Enabling Conversational Assistance for Point-of-Care Cardiac Ultrasound
图 1 · 摘自论文原文
  • 融合公开数据与真实手持设备采集,覆盖多种视角和图像质量。
  • 包含74,819组问答对,支持优化探头位置以获取标准左心室射血分数视图。
  • 提出轻量级多模态提示方法,性能达顶尖水平且参数量更少。

床旁经胸超声心动图(TTE)可在几乎任何临床场景中实现心脏评估,但其诊断价值受限于图像采集与解读所需的专业知识。视觉问答(VQA)为通过交互式临床辅助弥合这一差距提供了前景,但现有超声心动图VQA数据集规模有限、仅涵盖高质量图像且视图种类少。本文提出EchoVQA,首个大规模超声心动图VQA数据集,包含14,299张图像与74,819组问答对。数据集整合了公开来源(EchoNet-Dynamic、CAMUS)及我们使用两款手持探头(Lumify、Clarius)在床旁采集的图像,涵盖多样视图,并包含高质量与非理想图像。独特之处在于,该数据集包含用于指导用户优化探头位置以获得标准心尖四腔视图的问答对,从而提升左心室射血分数估算的准确性——这对新手操作者而言极具挑战。此外,我们提出一种基于多模态可学习提示的参数高效方法,在多数基准测试(包括EchoVQA)上达到当前最优性能,且可训练参数显著少于现有先进方法。

原文摘要 · Abstract (English)

Point-of-care transthoracic echocardiography (TTE) enables cardiac assessment in virtually any clinical setting, yet its diagnostic utility remains constrained by the expertise required for image acquisition and interpretation. Visual question answering (VQA) offers a promising paradigm for bridging this expertise gap through interactive clinical assistance, but existing echocardiography VQA datasets are limited in scale, restricted to high-quality images, and only cover a few views. We introduce EchoVQA, the first large-scale VQA dataset for echocardiography, comprising 14,299 images and 74,819 question-answer pairs. The dataset integrates public sources (EchoNet-Dynamic, CAMUS) with our own point-of-care acquisitions from two handheld probes (Lumify, Clarius), spanning diverse views and including both high-quality and suboptimal images. Uniquely, EchoVQA includes acquisition guidance questions to help users optimize transducer positioning toward a diagnostic apical 4-chamber view for left ventricular ejection fraction estimation -- a challenging task for novice operators in point-of-care settings. We further develop a parameter-efficient method based on multimodal learnable prompts achieving state-of-the-art performance on most benchmarks, including EchoVQA, with significantly less trainable parameters than existing state-of-the-art approaches.

超声心动图视觉问答医疗AI多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。