arXiv:2506.15304cs.CLcs.AI2025-06Conference of the …被引 5

用对比学习提升低资源语言识别准确率,尤其在跨领域数据上表现更好。

ConLID: Supervised Contrastive Learning for Low-Resource Language Identification

  • 采用监督对比学习,让模型学得跨领域通用的语言表征
  • 低资源语言在跨领域数据上识别准确率提升3.2个百分点
  • 适合需要高鲁棒性多语言处理的工业级文本预训练场景

语言识别(LID)是从小规模网页爬取中构建多语言大模型预训练语料的关键步骤。尽管许多研究致力于通过收集多样化训练数据来提升性能,但低资源语言(常仅限于单一领域数据,如圣经)仍表现不佳。为解决这一不平衡与偏差问题,我们提出一种新的监督对比学习(SCL)方法,用于学习低资源语言的领域不变表征。实验表明,该方法使低资源语言在跨领域数据上的识别性能提升3.2个百分点,同时保持对高资源语言的性能稳定。

原文摘要 · Abstract (English)

Language identification (LID) is a critical step in curating multilingual LLM pretraining corpora from web crawls. While many studies on LID model training focus on collecting diverse training data to improve performance, low-resource languages -- often limited to single-domain data, such as the Bible -- continue to perform poorly. To resolve these imbalance and bias issues, we propose a novel supervised contrastive learning (SCL) approach to learn domain-invariant representations for low-resource languages. We show that our approach improves LID performance on out-of-domain data for low-resource languages by 3.2 percentage points, while maintaining its performance for the high-resource languages.

语言识别对比学习低资源语言多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。