构建秘鲁医学考试数据集,评估大模型在西语医疗场景的答题能力。
PeruMedQA: Benchmarking Large Language Models (LLMs) on Peruvian Medical Exams -- Dataset Construction and Evaluation
- 构建8380道秘鲁专科医师考试题的多选问答数据集
- 微调后的medgemma-4b-it模型超越多数小模型,接近700亿参数模型表现
- 揭示西语医疗大模型性能差异,为拉美地区应用提供参考
背景:医疗大语言模型(LLMs)在解答医学考试方面表现出色,但其在西班牙语及拉丁美洲国家医学问题上的迁移能力尚未被探索。这一知识对医疗AI在拉丁美洲的应用至关重要。目标:构建秘鲁专科医师培训考试题目数据集;在该数据集上微调一个大语言模型;评估并比较原始模型与微调后模型的准确率表现。方法:我们构建了PeruMedQA数据集,包含8,380道多选题,覆盖12个专科领域(2018–2025年)。选取十种医学大模型,包括medgemma-4b-it和medgemma-27b-text-it,设计零样本任务提示进行作答。采用参数高效微调(PEFT)与低秩适配(LoRA)对medgemma-4b-it进行训练,使用除2025年外所有题目作为训练集。结果:medgemma-27b在各专科中表现最佳,精神科达最高分89.29%;在神经外科和放射科,OctoMed-7B分别以77.27%和77.38%、76.13%和77.39%略胜。多数参数少于100亿的模型正确率低于50%。微调后的medgemma-4b-it击败所有小于100亿参数的模型,并在多个专科中媲美700亿参数模型。结论:对于需基于西班牙语国家医学知识库且流行病学特征类似秘鲁的医疗AI应用,建议采用medgemma-27b-text-it。
原文摘要 · Abstract (English)
BACKGROUND: Medical large language models (LLMs) have demonstrated remarkable performance in answering medical examinations. However, the extent to which this high performance is transferable to medical questions in Spanish and from a Latin American country remains unexplored. This knowledge is crucial as LLM-based medical applications gain traction in Latin America. AIMS: To build a dataset of questions medical examinations taken by Peruvian physicians pursuing specialty training; to fine-tune a LLM on this dataset; to evaluate and compare the performance in terms of accuracy between vanilla LLMs and the fine-tuned LLM. METHODS: We curated PeruMedQA, a multiple-choice question-answering (MCQA) dataset containing 8,380 questions spanning 12 specialties (2018-2025). We selected ten medical LLMs, including medgemma-4b-it and medgemma-27b-text-it, and developed zero-shot task specific prompts to answer the questions. We employed parameter-efficient fine tuning (PEFT) and low-rand adaptation (LoRA) to fine-tune medgemma-4b-it utilizing all questions except those from 2025 (test set). RESULTS: Medgemma-27b showed the highest accuracy across all specialities, achieving the highest score of 89.29% in Psychiatry; yet, in two specialties, OctoMed-7B exhibited slight superiority: Neurosurgery with 77.27% and 77.38, respectively; and Radiology with 76.13% and 77.39%, respectively. Across specialties, most LLMs with <10 billion parameters exhibited <50% of correct answers. The fine-tuned version of medgemma-4b-it emerged victorious against all LLMs with <10 billion parameters and rivaled a LLM with 70 billion parameters across various examinations. CONCLUSIONS: For medical AI applications and research that require knowledge bases from Spanish-speaking countries and those exhibiting similar epidemiological profile to Peru's, interested parties should utilize medgemma-27b-text-it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。