arXiv:2504.09421cs.CLcs.AI2025-04被引 6

用2万份真实病历训练,提升大模型的疾病诊断推理能力

ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model

  • 基于2万份临床记录,融合多种训练策略增强诊断推理
  • 在中文任务中超越GPT-4o,英文表现接近GPT-4
  • 适合医疗AI研究者和临床辅助系统开发者参考

大型语言模型(LLM)在数学与编程领域展现出卓越的推理能力,但在临床诊断中的应用仍待深入。本文提出ClinicalGPT-R1,一个面向疾病诊断的增强推理通用大模型。该模型在包含20,000份真实临床记录的数据集上训练,采用多样化训练策略以提升诊断推理能力。为评估性能,我们构建了MedBench-Hard数据集,涵盖七大医学专科及代表性疾病。实验表明,ClinicalGPT-R1在中文诊断任务中优于GPT-4o,英文表现与GPT-4相当。该对比研究有效验证了其在疾病诊断任务中的优越性。相关资源见https://github.com/medfound/medfound。

原文摘要 · Abstract (English)

Recent advances in reasoning with large language models (LLMs)has shown remarkable reasoning capabilities in domains such as mathematics and coding, yet their application to clinical diagnosis remains underexplored. Here, we introduce ClinicalGPT-R1, a reasoning enhanced generalist large language model for disease diagnosis. Trained on a dataset of 20,000 real-world clinical records, ClinicalGPT-R1 leverages diverse training strategies to enhance diagnostic reasoning. To benchmark performance, we curated MedBench-Hard, a challenging dataset spanning seven major medical specialties and representative diseases. Experimental results demonstrate that ClinicalGPT-R1 outperforms GPT-4o in Chinese diagnostic tasks and achieves comparable performance to GPT-4 in English settings. This comparative study effectively validates the superior performance of ClinicalGPT-R1 in disease diagnosis tasks. Resources are available at https://github.com/medfound/medfound.

疾病诊断大模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。