构建首个阿拉伯语科学多选题数据集,评估大模型在理工科知识上的表现
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
- 创建覆盖多学科的阿拉伯语科学多选题数据集
- 现有大模型在该数据集上表现不佳,准确率普遍偏低
- 适合研究多语言模型、阿拉伯语NLP与教育AI的学者使用
大型语言模型(LLMs)不仅展现出生成类人文本的能力,还具备知识获取潜力,这要求我们超越传统的自然语言处理下游任务评估,全面考察模型的知识与推理能力。尽管已有大量基准测试,但多数集中于英语。鉴于许多大模型具备多语言能力,仅依赖英语评估不充分。为此,我们提出AraSTEM——一个用于评估大模型在理工科领域知识的阿拉伯语多选题数据集。该数据集涵盖多个学科、不同难度层级,要求模型深入理解科学类阿拉伯语才能取得高分。实验表明,公开可用的各类大小模型在此数据集上表现均不理想,凸显了对更本地化语言模型的需求。该数据集已开源至Hugging Face。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown remarkable capabilities, not only in generating human-like text, but also in acquiring knowledge. This highlights the need to go beyond the typical Natural Language Processing downstream benchmarks and asses the various aspects of LLMs including knowledge and reasoning. Numerous benchmarks have been developed to evaluate LLMs knowledge, but they predominantly focus on the English language. Given that many LLMs are multilingual, relying solely on benchmarking English knowledge is insufficient. To address this issue, we introduce AraSTEM, a new Arabic multiple-choice question dataset aimed at evaluating LLMs knowledge in STEM subjects. The dataset spans a range of topics at different levels which requires models to demonstrate a deep understanding of scientific Arabic in order to achieve high accuracy. Our findings show that publicly available models of varying sizes struggle with this dataset, and underscores the need for more localized language models. The dataset is freely accessible on Hugging Face.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。