arXiv:2603.09982cs.CLcs.AI2026-03中稿 · AbjadNLP Workshop,…被引 2

为阿拉伯语设计的长序列编码器,提升语言建模与下游任务表现

AraModernBERT: Transtokenized Initialization and Long-Context Encoder Modeling for Arabic

  • 采用跨字形分词初始化,显著提升阿拉伯语建模效果
  • 支持长达8192个标记的稳定长文本建模,语言模型性能更优
  • 适用于阿拉伯语理解、识别等任务,适配阿拉伯文字系统

编码器仅的Transformer模型在判别性NLP任务中仍广泛应用,但近期架构进展主要聚焦英语。本文提出AraModernBERT,将ModernBERT编码器架构适配至阿拉伯语,并研究跨字形分词嵌入初始化与原生长上下文建模(最长8,192个标记)的影响。结果表明,跨字形分词对阿拉伯语建模至关重要,相比非跨字形初始化,在掩码语言建模上取得显著提升。同时,AraModernBERT展现出稳定的长上下文建模能力,在延长序列长度下仍保持有效,内在语言建模性能优于基线。在阿拉伯语自然语言理解任务上的下游评估,包括推理、攻击性语言检测、问题相似度、命名实体识别,均证实其在判别性与序列标注任务中的强迁移能力。结果强调了将现代编码器架构适配阿拉伯语及其他阿拉伯字母书写系统时的关键实践考量。

原文摘要 · Abstract (English)

Encoder-only transformer models remain widely used for discriminative NLP tasks, yet recent architectural advances have largely focused on English. In this work, we present AraModernBERT, an adaptation of the ModernBERT encoder architecture to Arabic, and study the impact of transtokenized embedding initialization and native long-context modeling up to 8,192 tokens. We show that transtokenization is essential for Arabic language modeling, yielding dramatic improvements in masked language modeling performance compared to non-transtokenized initialization. We further demonstrate that AraModernBERT supports stable and effective long-context modeling, achieving improved intrinsic language modeling performance at extended sequence lengths. Downstream evaluations on Arabic natural language understanding tasks, including inference, offensive language detection, question-question similarity, and named entity recognition, confirm strong transfer to discriminative and sequence labeling settings. Our results highlight practical considerations for adapting modern encoder architectures to Arabic and other languages written in Arabic-derived scripts.

阿拉伯语长序列编码器语言建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。