arXiv:2604.10316cs.CL2026-04被引 1

对比多个大模型在医疗任务中的表现,发现专用模型更准,通用模型答题更快。

Comparative Analysis of Large Language Models in Healthcare

论文配图:Comparative Analysis of Large Language Models in Healthcare
图 1 · 摘自论文原文
  • 用医学数据集测试多个大模型的摘要与问答能力
  • 专用模型准确度高,通用模型答对率更高
  • 适合需要精准判断或快速作答的临床场景

大型语言模型(LLMs)正改变医疗领域的人工智能应用,因其能理解、生成和总结复杂医学文本。它们为临床医生、研究人员和患者提供支持,但在高风险临床环境中部署时,其准确性、可靠性与患者安全仍存疑。尽管近年关注度高,针对医疗应用的标准化基准评估仍有限。本研究旨在建立标准化的医学场景下大模型比较评估体系。我们评估了ChatGPT、LLaMA、Grok、Gemini和ChatDoctor等多个模型,在患者病历摘要与医学问答任务上的表现,使用MedMCQA、PubMedQA和Asclepius等开源数据集,通过语言学与任务特定指标综合评估。结果显示,领域专用模型如ChatDoctor在上下文可靠性方面表现优异,生成内容医学准确且语义一致;而通用模型如Grok和LLaMA在结构化问答任务中表现出更高的定量准确率。这表明不同任务下专用与通用模型具有互补优势。结论指出,大模型可有效辅助医疗决策,但其安全应用需遵守伦理标准、确保上下文准确,并保持人工监督。研究强调任务导向评估的重要性,以及在医疗流程中谨慎整合大模型的必要性。

原文摘要 · Abstract (English)

Background: Large Language Models (LLMs) are transforming artificial intelligence applications in healthcare due to their ability to understand, generate, and summarize complex medical text. They offer valuable support to clinicians, researchers, and patients, yet their deployment in high-stakes clinical environments raises critical concerns regarding accuracy, reliability, and patient safety. Despite substantial attention in recent years, standardized benchmarking of LLMs for medical applications has been limited. Objective: This study addresses the need for a standardized comparative evaluation of LLMs in medical settings. Method: We evaluate multiple models, including ChatGPT, LLaMA, Grok, Gemini, and ChatDoctor, on core medical tasks such as patient note summarization and medical question answering, using the open-access datasets, MedMCQA, PubMedQA, and Asclepius, and assess performance through a combination of linguistic and task-specific metrics. Results: The results indicate that domain-specific models, such as ChatDoctor, excel in contextual reliability, producing medically accurate and semantically aligned text, whereas general-purpose models like Grok and LLaMA perform better in structured question-answering tasks, demonstrating higher quantitative accuracy. This highlights the complementary strengths of domain-specific and general-purpose LLMs depending on the medical task. Conclusion: Our findings suggest that LLMs can meaningfully support medical professionals and enhance clinical decision-making; however, their safe and effective deployment requires adherence to ethical standards, contextual accuracy, and human oversight in relevant cases. These results underscore the importance of task-specific evaluation and cautious integration of LLMs into healthcare workflows.

大模型医疗AI评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。