评测大模型在医学诊断中的表现,发现部分模型已超医学生水平。
Performance of Large Language Models in Supporting Medical Diagnosis and Treatment
- 用葡萄牙医学专升考试数据集测试多款大模型性能
- 多个模型准确率超过医学生,且成本效益更优
- 适合临床辅助决策系统开发者参考
大型语言模型(LLMs)在医疗领域的集成具有显著潜力,可提升诊断准确率并支持治疗方案制定。这些基于AI的系统能分析海量数据,协助临床医生识别疾病、推荐治疗方案并预测患者预后。本研究评估了多种当代LLMs(含开源与闭源模型)在2024年葡萄牙医学专科准入国家考试(PNA)这一标准化医学知识测评上的表现。结果表明,各模型在准确率与成本效益方面存在显著差异,部分模型在该任务上的表现已超过医学生基准。我们基于准确率与成本的综合得分识别出领先模型,探讨了链式思维等推理方法的影响,并强调了LLMs作为复杂临床决策辅助工具的潜力。
原文摘要 · Abstract (English)
The integration of Large Language Models (LLMs) into healthcare holds significant potential to enhance diagnostic accuracy and support medical treatment planning. These AI-driven systems can analyze vast datasets, assisting clinicians in identifying diseases, recommending treatments, and predicting patient outcomes. This study evaluates the performance of a range of contemporary LLMs, including both open-source and closed-source models, on the 2024 Portuguese National Exam for medical specialty access (PNA), a standardized medical knowledge assessment. Our results highlight considerable variation in accuracy and cost-effectiveness, with several models demonstrating performance exceeding human benchmarks for medical students on this specific task. We identify leading models based on a combined score of accuracy and cost, discuss the implications of reasoning methodologies like Chain-of-Thought, and underscore the potential for LLMs to function as valuable complementary tools aiding medical professionals in complex clinical decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。