arXiv:2602.01714cs.CL2026-02被引 3

构建首个大规模阿拉伯语医学问答数据集,推动多语言医疗大模型发展。

MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark

  • 手动数字化阿拉伯语医学资料,构建跨19个专科的问答对数据集。
  • 涵盖五种难度级别,经专家与大模型双重评估,确保数据高质量。
  • 适合研究多语言医疗AI、临床部署及阿拉伯语NLP资源建设者。

阿拉伯语在自然语言处理研究中仍属严重资源匮乏的语言,尤其在医疗领域,受限于开源数据和基准测试资源稀缺。本文提出MedAraBench,一个大规模阿拉伯语医学多选题问答数据集,覆盖19个医学专科,分为五个难度等级。数据通过人工数字化阿拉伯语医疗专业人士撰写的学术材料构建,并经过严格预处理与划分训练/测试集。为评估质量,采用专家人工评价与“大模型作为评判者”双重框架。我们评估了八款先进开源及专有模型(如GPT-5、Gemini 2.0 Flash、Claude 4-Sonnet)的表现,结果表明当前模型仍需针对性优化。数据集与评测脚本已公开,旨在拓展医疗基准多样性,提升大模型多语言能力,助力临床应用落地。

原文摘要 · Abstract (English)

Arabic remains one of the most underrepresented languages in natural language processing research, particularly in medical applications, due to the limited availability of open-source data and benchmarks. The lack of resources hinders efforts to evaluate and advance the multilingual capabilities of Large Language Models (LLMs). In this paper, we introduce MedAraBench, a large-scale dataset consisting of Arabic multiple-choice question-answer pairs across various medical specialties. We constructed the dataset by manually digitizing a large repository of academic materials created by medical professionals in the Arabic-speaking region. We then conducted extensive preprocessing and split the dataset into training and test sets to support future research efforts in the area. To assess the quality of the data, we adopted two frameworks, namely expert human evaluation and LLM-as-a-judge. Our dataset is diverse and of high quality, spanning 19 specialties and five difficulty levels. For benchmarking purposes, we assessed the performance of eight state-of-the-art open-source and proprietary models, such as GPT-5, Gemini 2.0 Flash, and Claude 4-Sonnet. Our findings highlight the need for further domain-specific enhancements. We release the dataset and evaluation scripts to broaden the diversity of medical data benchmarks, expand the scope of evaluation suites for LLMs, and enhance the multilingual capabilities of models for deployment in clinical settings.

医学问答阿拉伯语大模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。