首个医学超声多任务评测基准,推动大模型在超声理解中的研究
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
- 构建涵盖15个解剖区域的7241例超声数据集,覆盖8类临床任务
- 23个主流视觉语言模型测试显示分类性能好,但空间推理与报告生成弱
- 适用于医疗AI研究者、超声诊断辅助系统开发者
超声是全球医疗中广泛应用的影像技术,但因其图像质量受操作者、噪声和解剖结构差异影响,解读仍具挑战。尽管大视觉语言模型(LVLMs)在自然与医疗领域展现出强大多模态能力,其在超声领域的表现仍缺乏系统评估。我们提出U2-BENCH,首个全面评估LVLMs在超声理解上的基准,涵盖分类、检测、回归与文本生成任务。该基准整合7,241例病例,覆盖15个解剖区域,定义8项临床相关任务,如诊断、视图识别、病灶定位、临床价值评估及报告生成,覆盖50种超声应用场景。我们评估了23个顶尖的开放与闭源LVLMs,包括通用与医学专用模型。结果表明,模型在图像级分类任务上表现良好,但在空间推理和临床语言生成方面仍存在显著挑战。U2-BENCH为评估与加速医学超声领域的大模型研究提供了严谨统一的测试平台。
原文摘要 · Abstract (English)
Ultrasound is a widely-used imaging modality critical to global healthcare, yet its interpretation remains challenging due to its varying image quality on operators, noises, and anatomical structures. Although large vision-language models (LVLMs) have demonstrated impressive multimodal capabilities across natural and medical domains, their performance on ultrasound remains largely unexplored. We introduce U2-BENCH, the first comprehensive benchmark to evaluate LVLMs on ultrasound understanding across classification, detection, regression, and text generation tasks. U2-BENCH aggregates 7,241 cases spanning 15 anatomical regions and defines 8 clinically inspired tasks, such as diagnosis, view recognition, lesion localization, clinical value estimation, and report generation, across 50 ultrasound application scenarios. We evaluate 23 state-of-the-art LVLMs, both open- and closed-source, general-purpose and medical-specific. Our results reveal strong performance on image-level classification, but persistent challenges in spatial reasoning and clinical language generation. U2-BENCH establishes a rigorous and unified testbed to assess and accelerate LVLM research in the uniquely multimodal domain of medical ultrasound imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。