arXiv:2505.21816cs.CL2025-05ACL被引 5

挑战阿拉伯语方言研究的四个常见假设,发现其过于简化现实。

Revisiting Common Assumptions about Arabic Dialects in NLP

  • 通过11国方言的多标签数据集,由母语者人工标注验证假设
  • 四条主流假设中部分不成立,方言差异远超区域划分
  • 为方言识别等任务提供更精准的理论基础,适合语言技术研究者

阿拉伯语拥有多种方言,不同方言间差异显著。现有自然语言处理文献普遍接受若干关于方言的假设(如“阿拉伯语方言可划分为明显区分的地区方言”),并在阿拉伯语方言识别等任务中应用。然而这些假设缺乏定量验证。本文识别出其中四个关键假设,并通过扩展和分析一个多标签数据集进行检验,该数据集包含11个国家级别的方言,每句文本均由对应方言母语者手动评估其有效性。分析表明,这四个假设过度简化了真实情况,部分并不总成立。这一发现可能阻碍阿拉伯语NLP任务的进一步发展。

原文摘要 · Abstract (English)

Arabic has diverse dialects, where one dialect can be substantially different from the others. In the NLP literature, some assumptions about these dialects are widely adopted (e.g., ``Arabic dialects can be grouped into distinguishable regional dialects") and are manifested in different computational tasks such as Arabic Dialect Identification (ADI). However, these assumptions are not quantitatively verified. We identify four of these assumptions and examine them by extending and analyzing a multi-label dataset, where the validity of each sentence in 11 different country-level dialects is manually assessed by speakers of these dialects. Our analysis indicates that the four assumptions oversimplify reality, and some of them are not always accurate. This in turn might be hindering further progress in different Arabic NLP tasks.

方言识别阿拉伯语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。