arXiv:2510.13430cs.CL2025-10综述被引 11

系统梳理阿拉伯语大模型评估基准,揭示短板并提出改进方向。

Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps

  • 按知识、任务、文化方言等四类组织40多个评估基准
  • 发现时序评估和多轮对话评测严重不足
  • 适合研究阿拉伯语NLP的学者参考

本综述首次系统分析阿拉伯语大模型评估基准,涵盖40多个跨自然语言处理任务、知识领域、文化理解与专项能力的评测集。提出四类分类体系:知识、自然语言任务、文化和方言、目标特定评估。分析显示基准多样性已有进展,但仍存在关键缺口:缺乏时序评估、多轮对话评测不足,以及翻译数据集中的文化偏差。探讨了原生构建、翻译和合成生成三种方法,比较其真实性、规模与成本的权衡。本工作为阿拉伯语NLP研究者提供全面参考,涵盖评估方法、可复现性标准与指标,并给出未来发展方向建议。

原文摘要 · Abstract (English)

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing benchmarks into four categories: Knowledge, NLP Tasks, Culture and Dialects, and Target-Specific evaluations. Our analysis reveals significant progress in benchmark diversity while identifying critical gaps: limited temporal evaluation, insufficient multi-turn dialogue assessment, and cultural misalignment in translated datasets. We examine three primary approaches: native collection, translation, and synthetic generation discussing their trade-offs regarding authenticity, scale, and cost. This work serves as a comprehensive reference for Arabic NLP researchers, providing insights into benchmark methodologies, reproducibility standards, and evaluation metrics while offering recommendations for future development.

阿拉伯语大模型评估基准测试NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。