让机器翻译更懂阿拉伯方言,可选地区和语体风格。
Context-Aware Dialectal Arabic Machine Translation with Interactive Region and Register Selection
- 用规则增强数据扩增3000句为5.7万句,覆盖8种方言
- 模型在8.19的BLEU下实现精准方言输出,优于主流系统
- 支持用户指定地区和语体,适合需要文化适配的场景
当前阿拉伯语机器翻译系统难以应对方言多样性,常将方言统一转为现代标准阿拉伯语(MSA),且缺乏对目标方言的用户控制。本文提出一种上下文感知、可调控的方言阿拉伯语翻译框架,显式建模区域与社会语言差异。核心技术是规则驱动的数据增强(RBDA)流程,将3000句种子语料扩展为5.7万句平衡的平行语料库,覆盖埃及、黎凡特、海湾等八种地区变体。通过在mT5-base模型上微调并加入轻量级元标签,实现对译文方言和语体的可控生成。自动评估与定性分析表明:高资源基线如NLLB(BLEU=13.75)虽得分高,但默认倾向MSA平均表达,方言特征弱;而本模型虽BLEU仅8.19,但输出更贴近目标方言。基于LLM辅助的文化真实性评估显示,本方法得分4.80/5,远超基线(1.0/5)。结果揭示标准指标对方言任务的局限性,呼吁建立更贴合阿拉伯语语言多样性的评估范式。
原文摘要 · Abstract (English)
Current Machine Translation (MT) systems for Arabic often struggle to account for dialectal diversity, frequently homogenizing dialectal inputs into Modern Standard Arabic (MSA) and offering limited user control over the target vernacular. In this work, we propose a context-aware and steerable framework for dialectal Arabic MT that explicitly models regional and sociolinguistic variation. Our primary technical contribution is a Rule-Based Data Augmentation (RBDA) pipeline that expands a 3,000-sentence seed corpus into a balanced 57,000-sentence parallel dataset, covering eight regional varieties eg., Egyptian, Levantine, Gulf, etc. By fine-tuning an mT5-base model conditioned on lightweight metadata tags, our approach enables controllable generation across dialects and social registers in the translation output. Through a combination of automatic evaluation and qualitative analysis, we observe an apparent accuracy-fidelity trade-off: high-resource baselines such as NLLB (No Language Left Behind) achieve higher aggregate BLEU scores (13.75) by defaulting toward the MSA mean, while exhibiting limited dialectal specificity. In contrast, our model achieves lower BLEU scores (8.19) but produces outputs that align more closely with the intended regional varieties. Supporting qualitative evaluation, including an LLM-assisted cultural authenticity analysis, suggests improved dialectal alignment compared to baseline systems (4.80/5 vs. 1.0/5). These findings highlight the limitations of standard MT metrics for dialect-sensitive tasks and motivate the need for evaluation practices that better reflect linguistic diversity in Arabic MT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。