针对阿拉伯字母语言,构建专用预训练模型提升文本分类效果。
The Role of Orthographic Consistency in Multilingual Embedding Models for Text Classification in Arabic-Script Languages
- 为每种阿拉伯字母语言定制预训练数据,聚焦书写规范特征。
- 在分类任务中比mBERT和XLM-RoBERTa高2~5个百分点。
- 适合研究阿拉伯字母语言、需高精度文本分类的场景。
在自然语言处理中,多语言模型如mBERT和XLM-RoBERTa虽覆盖广泛,但在共享阿拉伯字母但拼写规范与文化背景不同的语言(如库尔德语索拉尼、阿拉伯语、波斯语、乌尔都语)上表现不佳。本文提出阿拉伯字母RoBERTa(AS-RoBERTa)系列:四个基于RoBERTa的模型,分别在各自语言的大规模语料上进行预训练。通过聚焦语言特定的书写特征与统计规律,模型捕捉到通用模型忽略的模式。微调后,AS-RoBERTa变体在分类任务上比mBERT和XLM-RoBERTa提升2至5个百分点。消融实验表明,基于脚本的预训练是性能提升的关键。混淆矩阵分析揭示了共享脚本特性与领域内容对性能的影响。结果强调了对阿拉伯字母语言进行脚本感知专业化的重要性,并支持未来基于脚本与语言特异性的预训练策略。
原文摘要 · Abstract (English)
In natural language processing, multilingual models like mBERT and XLM-RoBERTa promise broad coverage but often struggle with languages that share a script yet differ in orthographic norms and cultural context. This issue is especially notable in Arabic-script languages such as Kurdish Sorani, Arabic, Persian, and Urdu. We introduce the Arabic Script RoBERTa (AS-RoBERTa) family: four RoBERTa-based models, each pre-trained on a large corpus tailored to its specific language. By focusing pre-training on language-specific script features and statistics, our models capture patterns overlooked by general-purpose models. When fine-tuned on classification tasks, AS-RoBERTa variants outperform mBERT and XLM-RoBERTa by 2 to 5 percentage points. An ablation study confirms that script-focused pre-training is central to these gains. Error analysis using confusion matrices shows how shared script traits and domain-specific content affect performance. Our results highlight the value of script-aware specialization for languages using the Arabic script and support further work on pre-training strategies rooted in script and language specificity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。