对比5个大模型在医疗问答中的零样本表现,发现大模型更优且部署需权衡效率。
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
- 用零样本评估法对比5个大模型在医学问答上的表现。
- 70B参数模型表现最佳,17B模型效率更高但性能略逊。
- 为医疗问答系统部署提供可复现的基准,适合资源受限场景。
近期大型语言模型(LLMs)在医疗领域受到广泛关注,尤其在构建医疗问答系统以提升资源匮乏地区的医疗可及性方面。本文对比了2024年4月至2025年8月间部署的五个LLMs在iCliniq数据集上的表现,该数据集包含38,000条来自不同专科的医学问答。所用模型包括Llama-3-8B-Instruct、Llama 3.2 3B、Llama 3.3 70B Instruct、Llama-4-Maverick-17B-128E-Instruct和GPT-5-mini。采用零样本评估方法,并使用BLEU与ROUGE指标进行无微调性能评估。结果表明,如Llama 3.3 70B Instruct等更大模型在性能上优于小模型,符合临床任务中观察到的缩放效应。值得注意的是,Llama-4-Maverick-17B展现出更具竞争力的表现,凸显了实际部署中计算效率与性能之间的权衡。这些发现与LLM向专业级医疗推理能力演进的趋势一致,反映了其在真实临床环境中支持问答系统的可行性日益提高。本研究旨在建立标准化评估基准,以最小化模型规模与计算资源消耗,最大化医疗自然语言处理应用的临床实用性。
原文摘要 · Abstract (English)
Recently, Large Language Models (LLMs) have gained significant traction in medical domain, especially in developing a QA systems to Medical QA systems for enhancing access to healthcare in low-resourced settings. This paper compares five LLMs deployed between April 2024 and August 2025 for medical QA, using the iCliniq dataset, containing 38,000 medical questions and answers of diverse specialties. Our models include Llama-3-8B-Instruct, Llama 3.2 3B, Llama 3.3 70B Instruct, Llama-4-Maverick-17B-128E-Instruct, and GPT-5-mini. We are using a zero-shot evaluation methodology and using BLEU and ROUGE metrics to evaluate performance without specialized fine-tuning. Our results show that larger models like Llama 3.3 70B Instruct outperform smaller models, consistent with observed scaling benefits in clinical tasks. It is notable that, Llama-4-Maverick-17B exhibited more competitive results, thus highlighting evasion efficiency trade-offs relevant for practical deployment. These findings align with advancements in LLM capabilities toward professional-level medical reasoning and reflect the increasing feasibility of LLM-supported QA systems in the real clinical environments. This benchmark aims to serve as a standardized setting for future study to minimize model size, computational resources and to maximize clinical utility in medical NLP applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。