首个阿拉伯语医学大模型评测基准,填补多语言医疗AI空白
MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks
- 构建涵盖7类任务的阿拉伯语医学数据集,含选择题、问答等
- 在GPT-4o等5个模型上测试,发现现有模型表现受限于语言适配
- 适合研究多语言医疗AI、公平性评估及低资源语言建模的学者
大型语言模型(LLMs)在医疗领域展现出巨大潜力,但其在阿拉伯语医学领域的有效性仍因缺乏高质量领域数据集和基准而未被充分探索。本研究提出MedArabiQ,一个包含七种阿拉伯语医学任务的新基准数据集,覆盖多个专科,包括多项选择题、填空题和医患问答。数据集基于过往医学考试和公开数据集构建,并引入多种改写方式以评估不同模型能力,如偏见缓解效果。我们对五种先进开源与专有模型(包括GPT-4o、Claude 3.5-Sonnet、Gemini 1.5)进行了全面评估。结果表明,亟需构建跨语言的高质量基准,以确保医疗LLM的公平部署与可扩展性。通过发布该基准与数据集,为未来研究提供了评估和提升多语言医疗大模型能力的基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated significant promise for various applications in healthcare. However, their efficacy in the Arabic medical domain remains unexplored due to the lack of high-quality domain-specific datasets and benchmarks. This study introduces MedArabiQ, a novel benchmark dataset consisting of seven Arabic medical tasks, covering multiple specialties and including multiple choice questions, fill-in-the-blank, and patient-doctor question answering. We first constructed the dataset using past medical exams and publicly available datasets. We then introduced different modifications to evaluate various LLM capabilities, including bias mitigation. We conducted an extensive evaluation with five state-of-the-art open-source and proprietary LLMs, including GPT-4o, Claude 3.5-Sonnet, and Gemini 1.5. Our findings highlight the need for the creation of new high-quality benchmarks that span different languages to ensure fair deployment and scalability of LLMs in healthcare. By establishing this benchmark and releasing the dataset, we provide a foundation for future research aimed at evaluating and enhancing the multilingual capabilities of LLMs for the equitable use of generative AI in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。