arXiv:2607.01908cs.CV2026-07

用大规模真实超声数据训练模型,简单方法实现高效临床理解。

Towards Real-World Ultrasound Understanding: Large Vision-Language Models from Multi-Image Examinations with Long-Form Reports

  • 以检查为单位对齐多图与未修剪报告,还原真实诊疗流程。
  • 1770万张图像、150万份检查数据,模型性能超越复杂设计的旧方法。
  • 适合医疗AI研究者和临床场景落地开发者参考。

大型视觉语言模型(LVLM)在众多医学影像任务中表现优异,但在超声领域的应用仍受限于其固有的复杂性和变异性。本文重新审视实现真实世界超声理解的核心需求:并非复杂的架构或繁琐的训练策略,而是数据规模与临床真实对齐。我们构建了一个包含150万份真实超声检查的大型数据集,涵盖1770万张图像,覆盖多器官,并配有未经编辑的临床报告。关键在于,数据按检查级别组织,将多幅图像与其对应报告对齐,以反映实际临床工作流。随后,我们使用低秩适应(LoRA)在该数据集上微调标准LVLM,无需任务特定修改。令人惊讶的是,这一简单方案已在多种超声理解任务中取得强效果,优于以往采用更复杂流水线的方法。此外,我们还进行了模型与数据扩展分析,揭示了规模在超声LVLM中的作用。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have achieved strong performance across many medical imaging tasks, yet their application to ultrasound remains limited due to its inherent complexity and variability. In this work, we revisit what is truly needed to enable real-world ultrasound understanding. Instead of introducing complex architectures or elaborate training strategies, we show that data scale and clinically faithful data alignment are the key factors. We construct a large-scale dataset of 1.5M real-world ultrasound examinations, containing 17.7M images, multi-organ coverage, and paired uncurated clinical reports. Crucially, we organize the data at the examination level, aligning multiple images with their corresponding reports to reflect real clinical workflows. We then fine-tune a standard LVLM using low-rank adaptation (LoRA) on this dataset without task-specific modifications. Surprisingly, this simple recipe already leads to strong performance across diverse ultrasound understanding tasks, outperforming prior methods designed with more complex pipelines. Beyond these results, we present model and data scaling analyses that provide insights into the role of scale in ultrasound LVLMs.

超声理解视觉语言模型医疗AI数据对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。