评测大模型在阿拉伯语医疗任务中的理解与推理能力,发现表现参差但有潜力。
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
- 用多模型投票提升多选题准确率,最高达77%。
- 开放问答中语义对齐最好达BERTScore 86.44%。
- 揭示当前阿拉伯语医疗大模型的优缺点,适合医疗AI研究者参考。
大型语言模型(LLMs)在阿拉伯语自然语言处理(NLP)应用中取得显著进展,但在阿拉伯语医疗NLP领域的有效性仍缺乏深入研究。本研究评估了前沿LLMs在阿拉伯语医疗任务中的知识表现与推理能力,基于MedArabiQ2025赛事中提出的AraHealthQA数据集开展基准测试。评估涵盖多项选择题(MCQs)、填空题及开放问答任务。结果显示,多选题任务中,融合Gemini Flash 2.5、Gemini Pro 2.5和GPT o3三个基座模型的多数投票方案表现最优,准确率达77%,在AraHealthQA 2025共享任务-子任务1中排名第一。开放问答任务中,部分模型在语义对齐上表现优异,最高BERTScore达到86.44%。整体表明,当前模型在阿拉伯语临床场景下具备潜力,但答案准确性与语义一致性之间存在明显差异。
原文摘要 · Abstract (English)
Recent progress in large language models (LLMs) has showcased impressive proficiency in numerous Arabic natural language processing (NLP) applications. Nevertheless, their effectiveness in Arabic medical NLP domains has received limited investigation. This research examines the degree to which state-of-the-art LLMs demonstrate and articulate healthcare knowledge in Arabic, assessing their capabilities across a varied array of Arabic medical tasks. We benchmark several LLMs using a medical dataset proposed in the Arabic NLP AraHealthQA challenge in MedArabiQ2025 track. Various base LLMs were assessed on their ability to accurately provide correct answers from existing choices in multiple-choice questions (MCQs) and fill-in-the-blank scenarios. Additionally, we evaluated the capacity of LLMs in answering open-ended questions aligned with expert answers. Our results reveal significant variations in correct answer prediction accuracy and low variations in semantic alignment of generated answers, highlighting both the potential and limitations of current LLMs in Arabic clinical contexts. Our analysis shows that for MCQs task, the proposed majority voting solution, leveraging three base models (Gemini Flash 2.5, Gemini Pro 2.5, and GPT o3), outperforms others, achieving up to 77% accuracy and securing first place overall in the Arahealthqa 2025 shared task-track 2 (sub-task 1) challenge. Moreover, for the open-ended questions task, several LLMs were able to demonstrate excellent performance in terms of semantic alignment and achieve a maximum BERTScore of 86.44%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。