arXiv:2506.07044cs.CLcs.AI2025-06被引 222

Lingshu 是一个专注医疗的多模态大模型,提升医学理解与推理能力。

Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

  • 构建覆盖影像与文本的医学数据集,融合精准标注与推理样本
  • 多阶段训练使模型在多项医疗任务中超越现有开源模型
  • 适合医疗AI研究者、临床辅助系统开发者使用

多模态大语言模型在通用视觉理解上表现优异,但在医疗领域受限于数据与任务差异。现有医学多模态模型存在三方面缺陷:(1)医学知识覆盖不足,仅限影像;(2)因数据质量差易产生幻觉;(3)缺乏复杂场景推理能力。为此,本文提出一套全面的数据筛选流程,从医学影像、大量医学文本及通用数据中高效获取丰富知识,并合成准确的图文描述、视觉问答与推理样本,构建涵盖广泛医学知识的多模态数据集。基于此,我们推出专用于医疗的多模态大模型 Lingshu,通过多阶段训练逐步嵌入医学专业能力。同时探索基于可验证奖励的强化学习以增强其推理能力。此外,开发 MedEvalKit 统一评估框架,整合主流多模态与文本医疗基准。在多模态问答、文本问答与医学报告生成三项基础任务上评估,结果表明 Lingshu 在多数任务中持续优于现有开源模型。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in understanding common visual elements, largely due to their large-scale datasets and advanced training strategies. However, their effectiveness in medical applications remains limited due to the inherent discrepancies between data and tasks in medical scenarios and those in the general domain. Concretely, existing medical MLLMs face the following critical limitations: (1) limited coverage of medical knowledge beyond imaging, (2) heightened susceptibility to hallucinations due to suboptimal data curation processes, (3) lack of reasoning capabilities tailored for complex medical scenarios. To address these challenges, we first propose a comprehensive data curation procedure that (1) efficiently acquires rich medical knowledge data not only from medical imaging but also from extensive medical texts and general-domain data; and (2) synthesizes accurate medical captions, visual question answering (VQA), and reasoning samples. As a result, we build a multimodal dataset enriched with extensive medical knowledge. Building on the curated data, we introduce our medical-specialized MLLM: Lingshu. Lingshu undergoes multi-stage training to embed medical expertise and enhance its task-solving capabilities progressively. Besides, we preliminarily explore the potential of applying reinforcement learning with verifiable rewards paradigm to enhance Lingshu's medical reasoning ability. Additionally, we develop MedEvalKit, a unified evaluation framework that consolidates leading multimodal and textual medical benchmarks for standardized, fair, and efficient model assessment. We evaluate the performance of Lingshu on three fundamental medical tasks, multimodal QA, text-based QA, and medical report generation. The results show that Lingshu consistently outperforms the existing open-source multimodal models on most tasks ...

医疗AI多模态大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。