arXiv:2409.16183cs.CV2024-09被引 5

面向真实医学影像的专家级多模态模型,显著提升诊断与报告生成能力。

Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

  • 基于810万张影像和25万对图文数据,定制放射科视觉语言架构。
  • 在三大真实模态上超越现有模型,人类评估与量化指标均领先。
  • 开源可复现,适合临床集成与医学AI研究者使用。

放射科是现代临床工作的重要且复杂的环节,涵盖多种任务。近期医疗视觉-语言(VL)基础模型在处理多模态信息方面展现出潜力,为多种放射科任务提供统一解决方案。然而,现有研究或在自然图像上预训练VL模型,或未充分融合视觉-语言架构与预训练,常忽视放射科图像及其文本上下文的独特多模态复杂性,且其在真实场景中的实用性尚未充分探索。本文提出RadFound,一个大规模、开源的放射科专用视觉-语言基础模型,基于超过810万张图像和25万对图像-文本配对数据进行训练,覆盖19个主要器官系统和10种成像模态。为实现专家级多模态感知与生成能力,RadFound引入增强的视觉编码器以捕捉图像内局部特征与图像间上下文信息,并采用专为放射科设计的统一跨模态学习框架。为全面评估模型能力,我们构建了基准测试RadVLBench,包含医学视觉-语言问答等解读任务,以及从图像描述到报告生成的文本生成任务,并提出人类评估框架。在包含三种代表性模态的真实世界基准测试中——二维图像(胸部X光)、多视角图像(乳腺钼靶)和三维图像(甲状腺CT扫描),RadFound在定量指标与人类评估中均显著优于其他VL基础模型。综上,RadFound的发展代表了放射科通用模型的进步,展示了其在临床工作流中广泛应用的潜力。

原文摘要 · Abstract (English)

Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in medicine have shown potential in processing multimodal information, offering a unified solution for various radiology tasks. However, existing studies either pre-trained VL models on natural data or did not fully integrate vision-language architecture and pretraining, often neglecting the unique multimodal complexity in radiology images and their textual contexts. Additionally, their practical applicability in real-world scenarios remains underexplored. Here, we present RadFound, a large and open-source vision-language foundation model tailored for radiology, that is trained on the most extensive dataset of over 8.1 million images and 250,000 image-text pairs, covering 19 major organ systems and 10 imaging modalities. To establish expert-level multimodal perception and generation capabilities, RadFound introduces an enhanced vision encoder to capture intra-image local features and inter-image contextual information, and a unified cross-modal learning design tailored to radiology. To fully assess the models' capability, we construct a benchmark, RadVLBench, including radiology interpretation tasks like medical vision-language question-answering, as well as text generation tasks ranging from captioning to report generation. We also propose a human evaluation framework. When evaluated on the real-world benchmark involving three representative modalities, 2D images (chest X-rays), multi-view images (mammograms), and 3D images (thyroid CT scans), RadFound significantly outperforms other VL foundation models on both quantitative metrics and human evaluation. In summary, the development of RadFound represents an advancement in radiology generalists, demonstrating broad applicability potential for integration into clinical workflows.

医学影像多模态基础模型放射科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。