arXiv:2601.13099cs.CL2026-01ACL被引 1

构建首个覆盖13国方言的阿拉伯语多轮对话翻译数据集,助力大模型理解真实口语。

Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs

  • 基于社区众包,采集107K轮跨方言对话,标注城市来源与说话人性别
  • 覆盖11个高影响力领域,包含医疗、教育等真实场景,支持细粒度方言建模
  • 适合作为多语言大模型训练与评估基准,尤其关注文化包容性与方言多样性

阿拉伯语是一种高度分化的语言,日常交流主要使用地区方言而非现代标准阿拉伯语(MSA)。然而,机器翻译系统对方言输入的泛化能力差,限制了数百万使用者的应用。我们提出Alexandria,一个大规模、社区驱动、人工翻译的数据集,旨在弥合这一差距。该数据集涵盖13个阿拉伯国家和11个高影响力领域,包括健康、教育、农业等。与以往资源不同,Alexandria通过城市来源元数据提供前所未有的细节,捕捉超越粗略区域标签的真实地方变体。数据集包含107,000条双语(英-方言)多轮对话,标注说话人与受话人性别配置,支持研究方言使用中的性别差异。该数据集既可用于训练,也可作为评估机器翻译和大语言模型(LLMs)在多样化阿拉伯语方言及次方言中表现的严格基准。我们的自动与人工评估揭示了当前阿拉伯语感知大模型在跨方言翻译中的能力边界,并暴露显著的持续挑战。Alexandria数据集、创建提示、翻译与修订指南以及评估代码已公开于GitHub:https://github.com/UBC-NLP/Alexandria

原文摘要 · Abstract (English)

Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than Modern Standard Arabic (MSA). Despite this, machine translation (MT) systems often generalize poorly to dialectal input, limiting their utility for millions of speakers. We introduce Alexandria, a large-scale, community-driven, human-translated dataset designed to bridge this gap. Alexandria covers 13 Arab countries and 11 high-impact domains, including health, education, and agriculture. Unlike previous resources, Alexandria provides unprecedented granularity by associating contributions with city-of-origin metadata, capturing authentic local varieties beyond coarse regional labels. The dataset consists of parallel English-Dialectal Arabic multi-turn conversational scenarios annotated with speaker-addressee gender configurations, enabling the study of gender-conditioned variation in dialectal use. Comprising 107K total turns, Alexandria serves as both a training resource and as a rigorous benchmark for evaluating MT and Large Language Models (LLMs). Our automatic and human evaluation benchmarks the current capabilities of Arabic-aware LLMs in translating across diverse Arabic dialects and sub-dialects while exposing significant persistent challenges. The Alexandria dataset, the creation prompts, the translation and revision guidelines, and the evaluation code are publicly available in the following repository: https://github.com/UBC-NLP/Alexandria

机器翻译阿拉伯语方言多轮对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。