评测大模型在波斯语医学问答中的表现,发现语言与领域适配至关重要。
PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark
- 构建20785题波斯语医学题集,覆盖23个专科,用于双语评估
- 通用模型如GPT-4.1在波斯语中达83.09%准确率,本地模型仅34.9%
- 文化临床上下文翻译丢失导致3-10%题目仅波斯语可答
大型语言模型(LLMs)在众多自然语言处理基准上表现卓越,常超越人类水平。然而,其在医疗等高风险领域,尤其是在低资源语言中的可靠性仍待深入探索。本文提出PersianMedQA,一个包含20,785道由专家验证的波斯语多选医学题的大规模数据集,源自伊朗14年国家医学考试,涵盖23个医学专科,旨在评估模型在波斯语和英语双语环境下的表现。我们对41个先进模型(包括通用、波斯语及医疗专用模型)在零样本与思维链(CoT)设置下进行评测。结果显示,封闭权重通用模型(如GPT-4.1)持续领先,波斯语中准确率达83.09%,英语中为80.7%;而波斯语专用模型(如Dorna)表现显著偏低(波斯语仅34.9%),普遍在指令遵循与领域推理上存在困难。我们还分析了翻译影响,发现尽管英语表现整体更高,但3-10%的问题因文化与临床上下文丢失,仅在波斯语中可正确解答。最后,证明模型大小不足以保证稳健表现,必须结合领域或语言适配。PersianMedQA为双语与文化相关医疗推理评估提供基础。数据集及双语医学词典已公开:https://huggingface.co/datasets/MohammadJRanjbar/PersianMedQA。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine, particularly in low-resource languages, remains underexplored. In this work, we introduce PersianMedQA, a large-scale dataset of 20,785 expert-validated multiple-choice Persian medical questions from 14 years of Iranian national medical exams, spanning 23 medical specialties and designed to evaluate LLMs in both Persian and English. We benchmark 41 state-of-the-art models, including general-purpose, Persian, and medical LLMs, in zero-shot and chain-of-thought (CoT) settings. Our results show that closed-weight general models (e.g., GPT-4.1) consistently outperform all other categories, achieving 83.09% accuracy in Persian and 80.7% in English, while Persian LLMs such as Dorna underperform significantly (e.g., 34.9% in Persian), often struggling with both instruction-following and domain reasoning. We also analyze the impact of translation, showing that while English performance is generally higher, 3-10% of questions can only be answered correctly in Persian due to cultural and clinical contextual cues that are lost in translation. Finally, we demonstrate that model size alone is insufficient for robust performance without strong domain or language adaptation. PersianMedQA provides a foundation for evaluating bilingual and culturally grounded medical reasoning in LLMs. The dataset, along with a bilingual medical dictionary, is available: https://huggingface.co/datasets/MohammadJRanjbar/PersianMedQA .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。