arXiv:2608.29990cs.CLcs.AI2026-09

首个针对沙特方言与文化能力的评分基准,揭示大模型在方言语用上的普遍短板。

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

  • 基于专家标注的31个真实场景提示,构建可量化的沙特方言评估体系。
  • 四款主流模型平均得分仅42.7%-53.1%,多数在特定提示中被扣分。
  • 错误主因是语境模糊而非编造事实,反映模型对语用细微差别的理解缺失。

大型语言模型在阿拉伯语市场应用日益广泛,但现有基准多偏重现代标准阿拉伯语(MSA)流畅性,忽视方言与文化语境能力。而日常交流以方言为主,且方言承载社会意义,仅以MSA评价无法捕捉其深层内涵。本文提出首个面向沙特方言的评分基准,包含31个专家撰写、涵盖习语、语用、词汇与文化嵌入现象的提示,每条均配有专家确立的真实答案。方法分为两阶段:先从真实答案中提炼出模型无关的原子化、互斥且穷尽(MECE)的正向标准;再对四种先进系统(Claude Opus 5、Gemini 3.7、GPT-5.6、Kimi K3)逐项评分,并对引入错误的项进行扣分。共完成124次模型-提示评估,记录466例错误,归类为九类。四模型宏观平均分集中在42.7%-53.1%区间,无一超过55%,且每个模型均有至少一条负分提示,证实沙特方言能力仍属未解难题。其中,模糊表述(Ambiguous Framing)占错误总数的37.3%,远超直接幻觉(11.2%),表明模型更常因扭曲语体、消解语用层次而失准,而非虚构内容。研究还发现一致性与上限间的权衡及模型特异性错误模式。完整提示集、真实答案与评分细则已公开,支持可复现的方言评估。

原文摘要 · Abstract (English)

Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks overwhelmingly reward Modern Standard Arabic (MSA) fluency while leaving dialectal and culturally grounded competence unmeasured. This gap is consequential: everyday Arabic is largely dialectal, and dialect encodes social meaning that MSA-centric evaluation cannot capture. We present a rubric-based benchmark for the Saudi dialect, comprising 31 expert-authored prompts spanning idiomatic, pragmatic, lexical, and culturally-embedded phenomena, each paired with an expert-established ground truth. Our methodology separates evaluation into a model-agnostic phase, in which atomic, MECE positive criteria are derived solely from the ground truth, and a model-specific phase, in which four state-of-the-art systems -- Claude Opus 5, Gemini 3.7, GPT-5.6, and Kimi K3 -- are scored against those criteria and penalised for errors they actively introduce. Across 124 model-prompt evaluations we catalogue 466 error instances under a nine-category taxonomy. The four systems cluster within a narrow macro-average band (42.7%-53.1%), with no model exceeding 55% and every model recording at least one negative-scoring prompt, confirming that Saudi dialectal competence remains broadly unsolved. Notably, Ambiguous Framing is the dominant failure mode (37.3% of errors) while outright Hallucination accounts for only 11.2%, indicating that models fail less by stating falsehoods than by distorting register and flattening pragmatic nuance. We further observe a consistency-versus-ceiling trade-off and model-distinctive error signatures. We release the full prompt set, ground truths, and scored rubrics to support reproducible dialectal evaluation.

方言评估文化理解LLM评测阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。