arXiv:2409.11404cs.CLcs.AI2024-09被引 66

构建首个阿拉伯方言与文化评估基准,提升大模型对本土语言文化的理解能力。

AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs

  • 用机器翻译加人工校对生成7个方言数据集,覆盖低资源方言。
  • 发现阿拉伯专用模型在方言任务上表现优于通用模型,但识别与生成仍存挑战。
  • 首次提出海湾、埃及、黎凡特地区文化细粒度评测,适合多语种研究者使用。

阿拉伯语因其丰富的方言多样性,在大语言模型中仍严重缺失,尤其在方言层面。本文通过机器翻译结合人工后编辑,构建了7个方言合成数据集,涵盖现代标准阿拉伯语(MSA)及地方变体。我们提出了AraDiCE——一个用于评估阿拉伯方言与文化能力的基准测试。该研究重点评估大模型在方言理解与生成方面的能力,尤其针对低资源方言。此外,首次引入针对海湾、埃及和黎凡特地区的细粒度文化认知评测,为大模型评估增加新维度。结果表明,尽管阿拉伯专用模型如Jais和AceGPT在方言任务上优于多语言模型,但在方言识别、生成与翻译方面仍面临显著挑战。本工作贡献约45,000条经人工校正的样本,以及文化评测基准,强调了定制化训练对捕捉阿拉伯语方言与文化细微差别的必要性。相关方言翻译模型与基准已公开发布于HuggingFace(https://huggingface.co/datasets/QCRI/AraDiCE)。

原文摘要 · Abstract (English)

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern Standard Arabic (MSA), created using Machine Translation (MT) combined with human post-editing. We present AraDiCE, a benchmark for Arabic Dialect and Cultural Evaluation. We evaluate LLMs on dialect comprehension and generation, focusing specifically on low-resource Arabic dialects. Additionally, we introduce the first-ever fine-grained benchmark designed to evaluate cultural awareness across the Gulf, Egypt, and Levant regions, providing a novel dimension to LLM evaluation. Our findings demonstrate that while Arabic-specific models like Jais and AceGPT outperform multilingual models on dialectal tasks, significant challenges persist in dialect identification, generation, and translation. This work contributes $\approx$45K post-edited samples, a cultural benchmark, and highlights the importance of tailored training to improve LLM performance in capturing the nuances of diverse Arabic dialects and cultural contexts. We have released the dialectal translation models and benchmarks developed in this study (https://huggingface.co/datasets/QCRI/AraDiCE).

阿拉伯语方言理解文化评测LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。