arXiv:2507.15850cs.CL2025-07被引 4

构建首个覆盖阿拉伯语科学、技术、代码的综合评测基准

3LM: Bridging Arabic, STEM, and Code through Benchmarking

  • 基于教材与练习题构建阿拉伯语科学问答数据集
  • 合成生成科学问题,提升数据多样性与覆盖度
  • 通过多轮人工校对翻译主流代码基准,保证质量

阿拉伯语是全球使用最广泛的语言之一,但针对其大语言模型(LLMs)的开发与评估仍相对有限。现有阿拉伯语评测多聚焦于语言、文化或宗教内容,而在真实世界应用中日益重要的科学、技术、工程与数学(STEM)及代码生成领域存在显著空白。为此,我们提出3LM,一套专为阿拉伯语设计的三个评测基准:第一个为源自阿拉伯语教材与教学练习的自然科学问答对;第二个为利用相同来源合成生成的科学问题;第三个聚焦代码生成,通过人工介入的多轮审校,将两个主流代码基准精准翻译成阿拉伯语。所有基准均公开发布,以推动阿拉伯语在这些关键但未充分代表领域的模型研究。

原文摘要 · Abstract (English)

Arabic is one of the most widely spoken languages in the world, yet efforts to develop and evaluate Large Language Models (LLMs) for Arabic remain relatively limited. Most existing Arabic benchmarks focus on linguistic, cultural, or religious content, leaving a significant gap in domains like STEM and code which are increasingly relevant for real-world LLM applications. To help bridge this gap, we present 3LM, a suite of three benchmarks designed specifically for Arabic. The first is a set of STEM-related question-answer pairs, naturally sourced from Arabic textbooks and educational worksheets. The second consists of synthetically generated STEM questions, created using the same sources. The third benchmark focuses on code generation, built through a careful translation of two widely used code benchmarks, incorporating a human-in-the-loop process with several rounds of review to ensure high-quality and faithful translations. We release all three benchmarks publicly to support the growth of Arabic LLM research in these essential but underrepresented areas.

阿拉伯语STEM评测代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。