arXiv:2502.18274cs.AIcs.CL2025-02被引 17

用专家思维路径训练医学大模型,提升诊断决策能力

Citrus: Leveraging Expert Cognitive Pathways in a Medical Language Model for Advanced Medical Decision Support

  • 模拟医生诊断时的思维过程,构建推理数据集
  • 在MedQA等基准上表现优于同规模模型
  • 开源诊断对话数据集,助力医疗AI研究

大型语言模型(LLM)在近年迅速发展,尤其在具备推理能力的模型中展现出广泛应用潜力。然而,在医疗领域,特别是在疾病推理任务中,其部署受限于专家级认知数据的获取难题。本文提出Citrus,一种通过模仿临床专家认知过程来弥合医学专长与AI推理差距的医学语言模型。该模型基于大规模模拟专家疾病推理数据进行训练,采用新颖方法精准捕捉临床决策路径,从而更真实地模拟复杂诊疗推理过程。为弥补医疗推理任务公开数据集的不足,研究团队发布了后期训练数据,包括自建的医学诊断对话数据集。在MedQA等权威基准上的评估显示,Citrus在医疗推理和语言理解任务中性能超越同规模其他模型。结果表明,Citrus有潜力显著提升医疗决策支持系统,提供更准确高效的临床辅助工具。

原文摘要 · Abstract (English)

Large language models (LLMs), particularly those with reasoning capabilities, have rapidly advanced in recent years, demonstrating significant potential across a wide range of applications. However, their deployment in healthcare, especially in disease reasoning tasks, is hindered by the challenge of acquiring expert-level cognitive data. In this paper, we introduce Citrus, a medical language model that bridges the gap between clinical expertise and AI reasoning by emulating the cognitive processes of medical experts. The model is trained on a large corpus of simulated expert disease reasoning data, synthesized using a novel approach that accurately captures the decision-making pathways of clinicians. This approach enables Citrus to better simulate the complex reasoning processes involved in diagnosing and treating medical conditions. To further address the lack of publicly available datasets for medical reasoning tasks, we release the last-stage training data, including a custom-built medical diagnostic dialogue dataset. This open-source contribution aims to support further research and development in the field. Evaluations using authoritative benchmarks such as MedQA, covering tasks in medical reasoning and language understanding, show that Citrus achieves superior performance compared to other models of similar size. These results highlight Citrus potential to significantly enhance medical decision support systems, providing a more accurate and efficient tool for clinical decision-making.

医疗AI大模型推理机制诊断支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。