arXiv:2508.17347cs.CLcs.AI2025-08EMNLP被引 2

提出阿拉伯语通用度得分,量化词汇跨方言使用范围。

The Arabic Generality Score: Another Dimension of Modeling Arabic Dialectness

  • 通过词对齐与历史编辑距离,计算词汇在方言间的通用程度。
  • 在多方言基准上优于现有顶尖方言识别系统。
  • 适合研究阿拉伯语方言连续体与词汇演变的学者使用。

阿拉伯语方言构成一个连续统一体,但自然语言处理模型常将其视为离散类别。近期工作通过阿拉伯方言度(ALDi)将方言度建模为连续变量,但该方法将复杂变异简化为单一维度。本文提出互补度量:阿拉伯语通用度得分(AGS),用于量化词汇在各方言中的广泛使用程度。我们设计了一套流程,结合词对齐、基于词源的编辑距离及平滑处理,对平行语料库进行词级AGS标注,并训练回归模型以预测上下文中的AGS。实验表明,该方法在多方言基准上超越了强基线,包括最先进的方言识别系统。AGS提供了一种可扩展且语言学合理的词汇通用性建模方式,丰富了阿拉伯语方言度表征。

原文摘要 · Abstract (English)

Arabic dialects form a diverse continuum, yet NLP models often treat them as discrete categories. Recent work addresses this issue by modeling dialectness as a continuous variable, notably through the Arabic Level of Dialectness (ALDi). However, ALDi reduces complex variation to a single dimension. We propose a complementary measure: the Arabic Generality Score (AGS), which quantifies how widely a word is used across dialects. We introduce a pipeline that combines word alignment, etymology-aware edit distance, and smoothing to annotate a parallel corpus with word-level AGS. A regression model is then trained to predict AGS in context. Our approach outperforms strong baselines, including state-of-the-art dialect ID systems, on a multi-dialect benchmark. AGS offers a scalable, linguistically grounded way to model lexical generality, enriching representations of Arabic dialectness.

方言建模词汇通用度阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。