构建首个阿拉伯语方言开放问答基准,评估模型跨文化理解能力
Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
- 将标准阿拉伯语多选题转为方言开放题,生成带思维链的标注数据
- 阿拉伯语方言上模型表现显著下降,文化知识存在明显短板
- 适合研究多语言、跨文化、方言适应性AI的学者与开发者
大型语言模型在回答日常问题中应用日益广泛,但在文化背景和方言内容上的表现仍不均衡。本文提出一种综合方法:(i) 将现代标准阿拉伯语(MSA)多选题(MCQ)翻译为英语及多种阿拉伯语方言;(ii) 转换为开放问答(OEQ)形式;(iii) 在MCQ与OEQ设置下评估零样本与微调后的多种大模型;(iv) 生成链式思维(CoT)推理过程以训练模型进行逐步推理。基于此方法,我们扩展了现有并行对齐多语言变体的问答数据集,据我们所知,这是首个此类数据集。实验涵盖开源与闭源模型。结果表明:(i) 模型在阿拉伯语方言上表现较差,暴露出文化相关与方言特异性知识的持续缺口;(ii) 以阿拉伯语为中心的模型在MCQ上表现良好,但在OEQ中困难重重;(iii) CoT提升了人工判断的正确率,但对n-gram指标影响不一。所构建的数据集将公开发布,以支持更具文化与语言包容性的评估研究。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal content remains uneven across languages. We propose a comprehensive method that (i) translates Modern Standard Arabic (MSA) multiple-choice questions (MCQs) into English and several Arabic dialects, (ii) converts them into open-ended questions (OEQs), (iii) benchmarks a range of zero-shot and fine-tuned LLMs under both MCQ and OEQ settings, and (iv) generates chain-of-thought (CoT) rationales to fine-tune models for step-by-step reasoning. Using this method, we extend an existing dataset in which QAs are parallelly aligned across multiple language varieties, making it, to our knowledge, the first of its kind. We conduct extensive experiments with both open and closed models. Our findings show that (i) models underperform on Arabic dialects, revealing persistent gaps in culturally grounded and dialect-specific knowledge; (ii) Arabic-centric models perform well on MCQs but struggle with OEQs; and (iii) CoT improves judged correctness while yielding mixed n-gram-based metrics. The developed dataset will be publicly released to support further research on culturally and linguistically inclusive evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。