arXiv:2506.01920cs.CL2025-06被引 11

构建阿拉伯语大模型评估新框架,填补文化理解与语言准确性空白

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation

  • 提出ADMD数据集,覆盖10大领域42子领域,聚焦深度文化与专业知识挑战
  • 评测5个主流模型,最佳表现仅30%准确率,数学与伊斯兰领域表现突出
  • 强调文化适配性对阿拉伯语模型评估的关键作用,适合研究者与开发者参考

本文针对阿拉伯语大模型评估中的关键缺陷,建立了全面的理论规范并提出新型评估框架。通过分析现有阿拉伯语评估数据集,发现其在语言准确性、文化契合度和方法严谨性方面存在严重问题。为此,我们构建了阿拉伯语深度微型数据集(ADMD),包含490个高难度问题,覆盖10大核心领域(42个子领域,见图1)。基于ADMD,我们评估了五款领先语言模型:GPT-4、Claude 3.5 Sonnet、Gemini Flash 1.5、CommandR 100B和Qwen-Max。结果表明,各模型在不同领域表现差异显著,尤其在需深层文化理解与专业背景的领域面临挑战。Claude 3.5 Sonnet总体准确率最高,达30%,在阿拉伯语、数学理论及伊斯兰相关领域表现相对优异。本工作为提升阿拉伯语大模型评估提供了理论基础与实践洞见,强调文化胜任力与技术能力同等重要。

原文摘要 · Abstract (English)

This paper addresses critical gaps in Arabic language model evaluation by establishing comprehensive theoretical guidelines and introducing a novel evaluation framework. We first analyze existing Arabic evaluation datasets, identifying significant issues in linguistic accuracy, cultural alignment, and methodological rigor. To address these limitations in LLMs, we present the Arabic Depth Mini Dataset (ADMD), a carefully curated collection of 490 challenging questions spanning ten major domains (42 sub-domains, see Figure 1. Using ADMD, we evaluate five leading language models: GPT-4, Claude 3.5 Sonnet, Gemini Flash 1.5, CommandR 100B, and Qwen-Max. Our results reveal significant variations in model performance across different domains, with particular challenges in areas requiring deep cultural understanding and specialized knowledge. Claude 3.5 Sonnet demonstrated the highest overall accuracy at 30\%, showing relative strength in mathematical theory in Arabic, Arabic language, and islamic domains. This work provides both theoretical foundations and practical insights for improving Arabic language model evaluation, emphasizing the importance of cultural competence alongside technical capabilities.

大模型评估阿拉伯语文化理解数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。