arXiv:2511.14295cs.CLcs.AI2025-11被引 1

评测大模型阿拉伯语能力,发现多数靠记忆而非真理解。

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

  • 人工标注150道题,覆盖语法、形态等五类语言能力
  • 35个模型测试显示,表面得分高但深层推理弱
  • 适合研究阿拉伯语大模型或语言理解的开发者

我们提出AraLingBench:一个完全由人工标注的基准,用于评估大语言模型(LLMs)在阿拉伯语语言能力方面的表现。该基准涵盖语法、形态、拼写、阅读理解与句法五个核心类别,包含150道专家设计的多项选择题,直接评估语言结构理解能力。对35个阿拉伯语及双语大模型的评估显示,当前模型虽在表层任务上表现良好,但在深层语法与句法推理方面仍存在明显短板。AraLingBench揭示了知识型基准高分与真实语言掌握之间的持续差距,表明许多模型依赖记忆或模式识别而非真正理解。通过分离并量化基础语言技能,AraLingBench为阿拉伯语大模型的发展提供了诊断框架。完整评估代码已公开于GitHub。

原文摘要 · Abstract (English)

We present AraLingBench: a fully human annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). The benchmark spans five core categories: grammar, morphology, spelling, reading comprehension, and syntax, through 150 expert-designed multiple choice questions that directly assess structural language understanding. Evaluating 35 Arabic and bilingual LLMs reveals that current models demonstrate strong surface level proficiency but struggle with deeper grammatical and syntactic reasoning. AraLingBench highlights a persistent gap between high scores on knowledge-based benchmarks and true linguistic mastery, showing that many models succeed through memorization or pattern recognition rather than authentic comprehension. By isolating and measuring fundamental linguistic skills, AraLingBench provides a diagnostic framework for developing Arabic LLMs. The full evaluation code is publicly available on GitHub.

阿拉伯语大模型评测语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。