arXiv:2604.03395cs.CL2026-04

QIMMA打造高质量阿拉伯语大模型评估基准,严控数据质量。

Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation

  • 用自动化+人工双校验,清理现有阿拉伯语评测集中的系统性错误
  • 构建超5.2万样本的多领域多任务评测集,主数据来自原生阿拉伯语内容
  • 开源推理结果与代码,支持复现与社区扩展

我们提出QIMMA,一个以系统性基准验证为核心的高质量阿拉伯语大模型排行榜。不同于直接整合现有资源,QIMMA通过结合自动化大模型判断与人工审核的多模型评估流程,在评测前识别并修复主流阿拉伯语基准中的系统性质量问题。最终形成一个经筛选的、涵盖多领域的多任务评测集,包含超过52,000个样本,主要基于原生阿拉伯语内容;代码评估任务为唯一例外,因其本质与语言无关。通过LightEval、EvalPlus实现透明化部署,并公开每样本推理输出,QIMMA成为可复现、可扩展的阿拉伯语NLP评估基础。

原文摘要 · Abstract (English)

We present QIMMA, a quality-assured Arabic LLM leaderboard that places systematic benchmark validation at its core. Rather than aggregating existing resources as-is, QIMMA applies a multi-model assessment pipeline combining automated LLM judgment with human review to surface and resolve systematic quality issues in well-established Arabic benchmarks before evaluation. The result is a curated, multi-domain, multi-task evaluation suite of over 52k samples, grounded predominantly in native Arabic content; code evaluation tasks are the sole exception, as they are inherently language-agnostic. Transparent implementation via LightEval, EvalPlus and public release of per-sample inference outputs make QIMMA a reproducible and community-extensible foundation for Arabic NLP evaluation.

阿拉伯语模型评估数据质量可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。