arXiv:2602.17054cs.CLcs.AI2026-02被引 2

ALPS挑战集测试阿拉伯语深层语义与语用理解,揭示模型在语法依赖上的严重缺陷。

ALPS: A Diagnostic Challenge Set for Arabic Linguistic & Pragmatic Reasoning

  • 构建531个专家设计的原生阿拉伯语诊断题,覆盖15任务47子任务。
  • 模型在带符号任务中语法错误率达36.5%,远高于语义理解误差。
  • 适合评估语言模型对阿拉伯语深层语法结构的真实理解能力。

当前阿拉伯语NLP基准多关注规模,常依赖合成或翻译数据,缺乏深层语言验证。本文提出ALPS(阿拉伯语语言与语用套件),一个原生、专家标注的诊断挑战集,聚焦深度语义与语用能力,弥补大规模基准的不足。该数据集包含531个精心设计的问题,覆盖15个任务和47个子任务,基于深厚的阿拉伯语语言学知识构建,确保文化真实性并消除翻译偏差。在23种不同模型(商业、开源及阿拉伯语专用)上评估,以单次人类表现(平均84.6%准确率)和专家裁定最优解(99.2%)为基准,发现模型虽流利但难以处理基础形态句法依赖:在依赖符号的任务中错误率达36.5%,显著高于组合语义任务。尽管顶级商业模型Gemini-3-flash达94.2%超越平均人类水平,但与阿拉伯语专用模型差距仍大,最佳本地模型Jais-2-70B仅达83.6%,未达人类表现。

原文摘要 · Abstract (English)

While recent Arabic NLP benchmarks focus on scale, they often rely on synthetic or translated data which may benefit from deeper linguistic verification. We introduce ALPS (Arabic Linguistic & Pragmatic Suite), a native, expert-curated diagnostic challenge set probing Deep Semantics and Pragmatics, capabilities that complement specialized large-scale benchmarks. While broad-coverage benchmarks prioritize scale and multi-task coverage, ALPS targets the depth of linguistic understanding through 531 rigorously crafted questions across 15 tasks and 47 subtasks. We developed the dataset with deep expertise in Arabic linguistics, guaranteeing cultural authenticity and eliminating translation artifacts. Evaluating 23 diverse models (commercial, open-source, and Arabic-native) against a single-pass human performance (avg. 84.6% accuracy) and an expert-adjudicated oracle (99.2%), we reveal a critical dissociation: models achieve high fluency but fail on fundamental morpho-syntactic dependencies, with elevated error rates on morpho-syntactic dependencies (36.5% across diacritics-reliant tasks) compared to compositional semantics. While top commercial models (Gemini-3-flash at 94.2%) surpass the average single human, a substantial gap persists between commercial giants and Arabic-native models, with the best Arabic-specific model (Jais-2-70B at 83.6%) approaching but not matching human performance.

阿拉伯语语义理解语言模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。