arXiv:2503.16463cs.AIcs.CL2025-03被引 3

通过临床数据训练提升大模型问诊能力,诊断准确率提升超30%。

Improving Interactive Diagnostic Ability of a Large Language Model Agent Through Clinical Experience Learning

  • 用350万份病历数据训练,聚焦初始问诊与诊断效率
  • 交互式问诊准确率提升30%以上,接近完整数据水平
  • 适合开发自主诊疗系统,尤其需高效信息收集的场景

大语言模型在医疗诊断中展现潜力,但其在需主动提问的交互式诊断中表现显著下降。本研究发现,问题主要出在初始诊断阶段的信息获取效率和初步判断形成,而非后续鉴别诊断。为此,我们提出一种可插拔增强(PPME)方法,基于中美医疗机构超过350万份电子病历,结合监督与强化学习,训练专用模型用于病史询问与初始诊断。实验表明,该方法使交互式诊断准确率相较基线提升30%以上,最终诊断水平接近使用完整临床数据的性能。结果表明该方法在构建自主诊断系统方面具有前景,但仍需进一步验证。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have shown promising results in medical diagnosis, with some studies indicating superior performance compared to human physicians in specific scenarios. However, the diagnostic capabilities of LLMs are often overestimated, as their performance significantly deteriorates in interactive diagnostic settings that require active information gathering. This study investigates the underlying mechanisms behind the performance degradation phenomenon and proposes a solution. We identified that the primary deficiency of LLMs lies in the initial diagnosis phase, particularly in information-gathering efficiency and initial diagnosis formation, rather than in the subsequent differential diagnosis phase. To address this limitation, we developed a plug-and-play method enhanced (PPME) LLM agent, leveraging over 3.5 million electronic medical records from Chinese and American healthcare facilities. Our approach integrates specialized models for initial disease diagnosis and inquiry into the history of the present illness, trained through supervised and reinforcement learning techniques. The experimental results indicate that the PPME LLM achieved over 30% improvement compared to baselines. The final diagnostic accuracy of the PPME LLM in interactive diagnostic scenarios approached levels comparable to those achieved using complete clinical data. These findings suggest a promising potential for developing autonomous diagnostic systems, although further validation studies are needed.

医疗AI大模型问诊优化临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。