arXiv:2501.02751eess.IVcs.CV2025-01被引 5

构建超声图像质量评估基准,测试大模型识别低质影像能力

Ultrasound-QBench: Can LLMs Aid in Quality Assessment of Ultrasound Imaging?

  • 设计三维度评估体系:分类、打分、对比,全面检验大模型表现
  • 涵盖超1.1万张真实超声图像,经专家标注分为高/中/低三类质量
  • 验证大模型具备初步识别超声伪影与质量的能力,助力临床诊断

随着超声检查数量激增,因操作者水平和成像条件差异导致的低质量图像日益增多,严重影响诊断准确性,甚至在关键情况下需重新检查。为帮助医生筛选高质量超声图像、保障诊断准确,我们提出Ultrasound-QBench,一个系统评估多模态大语言模型(MLLMs)在超声图像质量评估任务中表现的综合性基准。该基准包含两个数据集:来自不同来源的IVUSQA(7,709张图像)和CardiacUltraQA(3,863张图像),涵盖常见超声伪影,均由专业超声专家标注并分为高、中、低三个质量等级。为更有效评估MLLMs,我们将质量评估任务分解为三个维度:定性分类、定量评分和对比评估。对7个开源及1个专有MLLMs的评估表明,现有大模型已具备处理超声图像低级视觉任务的初步能力。我们希望此基准能推动研究界深入挖掘和提升大模型在医学影像任务中的潜力。

原文摘要 · Abstract (English)

With the dramatic upsurge in the volume of ultrasound examinations, low-quality ultrasound imaging has gradually increased due to variations in operator proficiency and imaging circumstances, imposing a severe burden on diagnosis accuracy and even entailing the risk of restarting the diagnosis in critical cases. To assist clinicians in selecting high-quality ultrasound images and ensuring accurate diagnoses, we introduce Ultrasound-QBench, a comprehensive benchmark that systematically evaluates multimodal large language models (MLLMs) on quality assessment tasks of ultrasound images. Ultrasound-QBench establishes two datasets collected from diverse sources: IVUSQA, consisting of 7,709 images, and CardiacUltraQA, containing 3,863 images. These images encompassing common ultrasound imaging artifacts are annotated by professional ultrasound experts and classified into three quality levels: high, medium, and low. To better evaluate MLLMs, we decompose the quality assessment task into three dimensionalities: qualitative classification, quantitative scoring, and comparative assessment. The evaluation of 7 open-source MLLMs as well as 1 proprietary MLLMs demonstrates that MLLMs possess preliminary capabilities for low-level visual tasks in ultrasound image quality classification. We hope this benchmark will inspire the research community to delve deeper into uncovering and enhancing the untapped potential of MLLMs for medical imaging tasks.

超声影像大模型质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。