arXiv:2601.22439cs.CL2026-01中稿 · LoResLM 2025被引 1

通过自适应负采样缓解低资源语言词元被忽略的问题

Stop Jostling: Adaptive Negative Sampling Reduces the Marginalization of Low-Resource Language Tokens by Cross-Entropy Loss

  • 设计阈值机制,动态减少罕见词元的边际化影响
  • 在字符级模型上提升低资源语言验证集性能
  • 适合关注少数语言表示与公平性的研究者

神经语言模型因训练数据有限,常难以有效学习低资源语言。本文针对训练过程中罕见词元受过度边际化影响的问题,提出一种阈值化方法,降低其负面影响,使稀有词元获得更有效的语义对齐。在字符级语言模型上的实验表明,该方法显著提升了低资源语言的验证表现。这是首个将负采样应用于缓解罕见词元边际化的研究,为增强少数语言建模能力提供了新思路。

原文摘要 · Abstract (English)

Neural language models often struggle with low-resource languages due to the limited availability of training data, making tokens from these languages rare in the training set. This paper addresses a specific challenge during training: rare tokens are disproportionately affected by marginalization, which prevents them from learning effectively. We propose a thresholding technique that reduces the impact of this marginalization, allowing rare tokens to benefit from more meaningful alignment. Through experiments with a character-level language model, we demonstrate that this method significantly improves performance on low-resource language validation data. This work is the first to show how negative sampling can be applied to improve the representation of rare tokens by limiting the harmful influence of excessive marginalization, offering a new approach to enhancing language model performance for underrepresented languages.

语言模型低资源语言负采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。