arXiv:2511.08614cs.CL2025-11

用多个大模型集成提升急诊诊断准确率,最高达85%。

A Super-Learner with Large Language Models for Medical Emergency Advising

  • 将五个主流大模型整合为超级学习器,通过元学习融合各自优势。
  • 集成系统诊断准确率达70%,单个模型最高达85%。
  • 适合急诊医疗辅助场景,尤其在医生资源紧张时应用价值高。

医疗决策支持与咨询系统对急诊医生快速准确评估患者状况并做出诊断至关重要。近年来,人工智能(AI)在医疗领域迅速发展,大型语言模型(LLMs)被广泛应用于医疗决策支持系统中。我们研究了五种主流大模型在真实急诊病例中的表现,结果显示其对急症诊断的准确率介于58%至65%之间,显著高于人类医生的报告准确率。为此,我们构建了由Gemini、Llama、Grok、GPT和Claude组成的超级学习器MEDAS(Medical Emergency Diagnostic Advising System)。该系统采用基础元学习器,整体诊断准确率达70%,且其中至少一个集成模型可达到85%的正确诊断率。超级学习器通过元学习机制,识别并利用各模型在不同任务上的能力差异,实现多模型集体知识的协同增益。研究结果表明,基于元学习的聚合诊断准确率超越任一单独模型,证明该系统能有效整合各模型所训练的医学数据集知识。

原文摘要 · Abstract (English)

Medical decision-support and advising systems are critical for emergency physicians to quickly and accurately assess patients' conditions and make diagnosis. Artificial Intelligence (AI) has emerged as a transformative force in healthcare in recent years and Large Language Models (LLMs) have been employed in various fields of medical decision-support systems. We studied responses of a group of different LLMs to real cases in emergency medicine. The results of our study on five most renown LLMs showed significant differences in capabilities of Large Language Models for diagnostics acute diseases in medical emergencies with accuracy ranging between 58% and 65%. This accuracy significantly exceeds the reported accuracy of human doctors. We built a super-learner MEDAS (Medical Emergency Diagnostic Advising System) of five major LLMs - Gemini, Llama, Grok, GPT, and Claude). The super-learner produces higher diagnostic accuracy, 70%, even with a quite basic meta-learner. However, at least one of the integrated LLMs in the same super-learner produces 85% correct diagnoses. The super-learner integrates a cluster of LLMs using a meta-learner capable of learning different capabilities of each LLM to leverage diagnostic accuracy of the model by collective capabilities of all LLMs in the cluster. The results of our study showed that aggregated diagnostic accuracy provided by a meta-learning approach exceeds that of any individual LLM, suggesting that the super-learner can take advantage of the combined knowledge of the medical datasets used to train the group of LLMs.

急诊诊断大模型集成医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。