针对阿尔及利亚方言的复杂形态,提出字符级语言模型chDzDT。
chDzDT: Word-level morphology-aware language model for Algerian social media text
- 基于字符级建模,绕过分词限制,捕捉方言形态特征。
- 在多源数据上训练,覆盖多种语言和书写系统。
- 适合低资源方言处理,尤其对形态丰富的语言有优势。
预训练语言模型(PLMs)通过提供上下文敏感的文本表示,显著推动了自然语言处理的发展。然而,阿尔及利亚方言仍缺乏专门模型,其复杂形态、频繁混用语言、多重书写系统及外来词汇影响,使传统词或子词级别的方法难以有效处理。为此,我们提出chDzDT,一种专为阿尔及利亚方言设计的字符级预训练语言模型。该模型不依赖词序列,而是以孤立单词为单位进行训练,从而稳健编码形态模式,不受分词边界或标准拼写的影响。训练语料来自YouTube评论、法语、英语和柏柏尔语维基百科以及Tatoeba项目,涵盖多种语言与书写形式,训练负荷较大。贡献包括:(i) 利用YouTube评论对阿尔及利亚方言进行细致形态分析;(ii) 构建多语言阿尔及利亚词典数据集;(iii) 开发并全面评估字符级PLM作为下游任务的形态感知编码器。结果表明,字符级建模在形态丰富且资源匮乏的方言中具有潜力,为更包容、灵活的NLP系统奠定基础。
原文摘要 · Abstract (English)
Pre-trained language models (PLMs) have substantially advanced natural language processing by providing context-sensitive text representations. However, the Algerian dialect remains under-represented, with few dedicated models available. Processing this dialect is challenging due to its complex morphology, frequent code-switching, multiple scripts, and strong lexical influences from other languages. These characteristics complicate tokenization and reduce the effectiveness of conventional word- or subword-level approaches. To address this gap, we introduce chDzDT, a character-level pre-trained language model tailored for Algerian morphology. Unlike conventional PLMs that rely on token sequences, chDzDT is trained on isolated words. This design allows the model to encode morphological patterns robustly, without depending on token boundaries or standardized orthography. The training corpus draws from diverse sources, including YouTube comments, French, English, and Berber Wikipedia, as well as the Tatoeba project. It covers multiple scripts and linguistic varieties, resulting in a substantial pre-training workload. Our contributions are threefold: (i) a detailed morphological analysis of Algerian dialect using YouTube comments; (ii) the construction of a multilingual Algerian lexicon dataset; and (iii) the development and extensive evaluation of a character-level PLM as a morphology-focused encoder for downstream tasks. The proposed approach demonstrates the potential of character-level modeling for morphologically rich, low-resource dialects and lays a foundation for more inclusive and adaptable NLP systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。