arXiv:2607.00139cs.CL2026-07被引 1

评测大模型对阿拉伯文化语言知识的掌握,发现其在伊拉克方言上表现明显更差。

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

论文配图:Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth
图 1 · 摘自论文原文
  • 用母语专家标注的103组提示-评分对,构建跨社区评估框架
  • 模型在埃及方言任务上表现优于伊拉克方言,但专家打分存在系统性宽松偏差
  • 自动评分易受文化推理误导,需模拟母语者判断而非仅查词

人工专家评估成本高,尤其在阿拉伯社会语言学这类需深厚文化背景的领域。本文针对埃及和伊拉克阿拉伯语两种方言,构建跨社区评估框架,由母语专家创作并评分103组提示-评分对(70个埃及、33个伊拉克;53项文化、50项语言),采用带惩罚权重的评分标准区分正面内容要求与具体错误。三款前沿大模型在302个独立提示-响应对上接受专家评分,五款模型作为自动化评判者进行自检。双指标方案结合平均绝对偏差(MAD)与有符号均值误差,分离出方向性评分偏倚与对称噪声。1307次评判显示:GPT-5.4是最可靠的评判者(MADj = 10.21pp,有符号误差 = -1.12%);五名评判者中有四名系统性偏松(+2.01%至+6.56%);所有评判者对文化类任务评分更难(MAD差距1.83-4.78pp);模型在埃及提示上显著优于伊拉克提示。但由于伊拉克与埃及专家评分宽松程度不同,无法单归因于模型知识。因此强调不假设评分者一致性。所有样本中,隐含文化推理——即模型需模拟母语者判断而非仅依赖词汇验证——成为自动评分的主要失败模式。

原文摘要 · Abstract (English)

The cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, high-stakes domains. This is particularly acute for Arabic sociolinguistic knowledge: credible grading requires not only linguistic fluency but deep cultural familiarity that cannot be approximated by surface-level metrics. We address this with a cross-evaluation framework instantiated on two underrepresented Arabic dialect communities: Egyptian and Iraqi Arabic. We contribute 103 validated prompt-rubric pairs (70 Egyptian, 33 Iraqi; 53 Cultural, 50 Linguistic), authored and graded by native-speaker SMEs using penalty-weighted rubrics distinguishing positive content requirements from answer-specific negative error criteria. Three frontier LLMs serve as target models (graded by human SMEs across 302 unique prompt-response pairs), while five frontier LLMs serve as automated judges enforcing a provider-level self-evaluation guard. A dual-metric scheme combining Mean Absolute Deviation (MAD) with Signed Mean Error separates directional grading bias from symmetric noise. Across 1,307 judge evaluations: GPT-5.4 is the most reliable judge (MADj = 10.21 pp, Signed Error = -1.12%); four of five judges show systematic leniency (+2.01% to +6.56%); Cultural tasks are harder to grade than Linguistic tasks for all judges (MAD gap 1.83-4.78 pp); and models substantially outperform on Egyptian prompts compared to Iraqi prompts. However, given leniency differences between Iraqi and Egyptian SMEs, we cannot solely attribute this gap to model knowledge. We therefore emphasize findings that do not assume identical leniency across human graders. Across all samples, implicit cultural reasoning -- requiring models to simulate native-speaker judgment rather than rely on lexical verification -- emerges as the primary failure mode for automated grading across all judge models.

大模型评测阿拉伯语文化推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。