首个胎儿超声多模态评估基准,助力AI提升产前筛查效率
FETAL-GAUGE: A Benchmark for Assessing Vision-Language Models in Fetal Ultrasound
- 构建涵盖4.2万张图像、9.3万个问答对的胎儿超声视觉问答基准
- 顶尖模型在多项任务上准确率仅55%,远低于临床需求
- 适用于医学AI研究者与产科医疗智能化开发者
产前超声检查需求激增导致训练有素的超声技师全球短缺,阻碍胎儿健康监测。深度学习有望提升技师效率并辅助新人培训。视觉语言模型(VLMs)因其能联合处理图像与文本,在单一框架内完成多项临床任务而前景广阔。然而,由于该模态存在操作依赖性强、数据公开难等挑战,目前尚无标准化的胎儿超声评估基准。为此,我们提出Fetal-Gauge,首个专用于评估VLMs在胎儿超声中表现的大型视觉问答基准。该基准包含超过42,000张图像和93,000个问题-答案对,覆盖解剖切面识别、结构定位、胎儿方位判断、视图符合度及临床诊断等任务。我们系统评估了多种先进VLMs(包括通用与医学专用模型),发现最佳模型准确率仅为55%,远未达临床要求。分析揭示当前VLMs在胎儿超声解读中的关键局限,凸显亟需领域适配架构与专用训练方法。Fetal-Gauge为推进产前护理中的多模态深度学习奠定坚实基础,并为应对全球医疗可及性挑战提供路径。基准将在论文录用后公开。
原文摘要 · Abstract (English)
The growing demand for prenatal ultrasound imaging has intensified a global shortage of trained sonographers, creating barriers to essential fetal health monitoring. Deep learning has the potential to enhance sonographers' efficiency and support the training of new practitioners. Vision-Language Models (VLMs) are particularly promising for ultrasound interpretation, as they can jointly process images and text to perform multiple clinical tasks within a single framework. However, despite the expansion of VLMs, no standardized benchmark exists to evaluate their performance in fetal ultrasound imaging. This gap is primarily due to the modality's challenging nature, operator dependency, and the limited public availability of datasets. To address this gap, we present Fetal-Gauge, the first and largest visual question answering benchmark specifically designed to evaluate VLMs across various fetal ultrasound tasks. Our benchmark comprises over 42,000 images and 93,000 question-answer pairs, spanning anatomical plane identification, visual grounding of anatomical structures, fetal orientation assessment, clinical view conformity, and clinical diagnosis. We systematically evaluate several state-of-the-art VLMs, including general-purpose and medical-specific models, and reveal a substantial performance gap: the best-performing model achieves only 55\% accuracy, far below clinical requirements. Our analysis identifies critical limitations of current VLMs in fetal ultrasound interpretation, highlighting the urgent need for domain-adapted architectures and specialized training approaches. Fetal-Gauge establishes a rigorous foundation for advancing multimodal deep learning in prenatal care and provides a pathway toward addressing global healthcare accessibility challenges. Our benchmark will be publicly available once the paper gets accepted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。