arXiv:2606.22478cs.CL2026-06

针对罗马乌尔都语词汇碎片化问题,提出稳定嵌入空间的词表扩展方法

ROMEVA: Geometry-Preserving Vocabulary Expansion for Roman Urdu Language Models

  • 用子词平均初始化+主成分引导锚点损失,减少词表扩展时嵌入偏移
  • 在3.6万条评论数据上扩增500个碎片化词条,保持嵌入空间稳定性最优
  • 发现下游任务性能不依赖嵌入稳定,适配性强比保持原空间更重要

多语言模型如mBERT广泛用于低资源自然语言处理,但其在形态不一致语言(如罗马乌尔都语)中的适配仍不充分。罗马乌尔都语拼写变体导致严重子词碎片化,平均每个词元产生1.50个子词。本文提出ROMEVA(罗马乌尔都语嵌入保持词表适配),结合子词平均初始化与主成分分析引导的锚点损失,以稳定词表扩展过程中的嵌入空间。基于36,130条罗马乌尔都语评论语料,向mBERT添加500个高度碎片化词元,并对比朴素微调、子词感知微调与ROMEVA的效果。结果表明,ROMEVA最有效保持预训练嵌入空间,而朴素微调在下游情感分类任务中表现最佳。这一发现揭示嵌入稳定性与下游性能间的脱节,提示在形态不一致语言中,更强的适应性可能优于严格的嵌入保持。

原文摘要 · Abstract (English)

Multilingual Language Models like mBERT are widely used for low-resource NLP, yet their adaptation to morphologically inconsistent languages such as Roman Urdu remains underexplored. Roman Urdu spelling variation causes severe sub-word fragmentation, averaging 1.50 sub-words per token. We propose \textit{ROMEVA} (Roman Urdu Embedding-preserving Vocabulary Adaptation), which combines sub-word-average initialization and a PCA-guided anchor loss to stabilize embeddings during vocabulary expansion. Using a 36,130-comment Roman Urdu corpus, we add 500 highly fragmented tokens to mBERT and compare naive fine-tuning, sub-word-aware fine-tuning, and \textit{ROMEVA}. While \textit{ROMEVA} most effectively preserves the pretrained embedding space, naive fine-tuning achieves the strongest downstream sentiment classification performance. These findings reveal a disconnect between embedding stability and downstream performance, suggesting that stronger adaptation may be preferable to strict embedding preservation in morphologically inconsistent languages.

词表扩展罗马乌尔都语嵌入保持低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。