arXiv:2501.05464cs.CLcs.AI2025-01被引 47

用病例生成提升大模型医疗问答准确率

LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language Models

  • 通过多智能体架构生成相似病例增强推理
  • 零样本下准确率和F1提升7个百分点
  • 适合医疗AI研发与临床辅助系统应用

精准高效的问答系统对高质量患者护理至关重要。尽管大语言模型在多个领域取得显著进展,但在医疗问答中仍面临理解专业术语和复杂推理的挑战,限制了其在关键医疗场景中的应用。为此,我们提出一种新方法,在多智能体医疗问答系统中引入相似病例生成。具体而言,采用Llama3.1:70B这一前沿大模型,在零样本学习条件下构建多智能体架构以提升在MedQA数据集上的表现。该方法充分利用模型固有的医学知识与推理能力,无需额外训练数据。实验结果表明,相较于现有基准模型,在多种医疗问答任务中准确率与F1分数均提升7%。此外,我们还评估了模型在处理复杂医疗问题时的可解释性与可靠性。本研究不仅为医疗问答提供可靠解决方案,也为大模型在医疗领域的更广泛应用奠定基础。

原文摘要 · Abstract (English)

Accurate and efficient question-answering systems are essential for delivering high-quality patient care in the medical field. While Large Language Models (LLMs) have made remarkable strides across various domains, they continue to face significant challenges in medical question answering, particularly in understanding domain-specific terminologies and performing complex reasoning. These limitations undermine their effectiveness in critical medical applications. To address these issues, we propose a novel approach incorporating similar case generation within a multi-agent medical question-answering (MedQA) system. Specifically, we leverage the Llama3.1:70B model, a state-of-the-art LLM, in a multi-agent architecture to enhance performance on the MedQA dataset using zero-shot learning. Our method capitalizes on the model's inherent medical knowledge and reasoning capabilities, eliminating the need for additional training data. Experimental results show substantial performance gains over existing benchmark models, with improvements of 7% in both accuracy and F1-score across various medical QA tasks. Furthermore, we examine the model's interpretability and reliability in addressing complex medical queries. This research not only offers a robust solution for medical question answering but also establishes a foundation for broader applications of LLMs in the medical domain.

医疗问答大模型多智能体零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。