arXiv:2511.10229cs.CL2025-11中稿 · ed被引 3

通过语言可分性筛选数据,提升多语言大模型训练效果

LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning

  • 基于语言在模型表征中的可区分度筛选训练数据
  • 在22种语言上提升低资源语言和理解类任务表现
  • 适合关注多语言模型训练数据优化的研究者

联合多语言指令微调是提升大语言模型多语言指令遵循能力与下游性能的常用方法,但其效果高度依赖训练数据的组成与选择。现有方法通常基于文本质量、多样性或任务相关性等特征,却忽略了多语言数据内在的语言结构。本文提出LangGPS,一种轻量级两阶段预筛选框架,基于语言可分性——即不同语言样本在模型表示空间中可区分的程度——进行数据筛选。首先根据可分性分数过滤数据,再结合已有筛选方法进一步优化。在六个基准测试和22种语言上的实验表明,将LangGPS应用于现有筛选方法可显著提升其在多语言训练中的有效性与泛化能力,尤其在理解任务和低资源语言上表现突出。进一步分析发现,高可分性样本有助于形成清晰的语言边界并加速模型适应,而低可分性样本则充当跨语言对齐的桥梁。此外,语言可分性还可作为多语言课程学习的有效信号,交错不同可分性水平的样本能带来稳定且通用的提升。本工作为多语言语境下数据效用提供了新视角,助力更贴近语言规律的大模型发展。

原文摘要 · Abstract (English)

Joint multilingual instruction tuning is a widely adopted approach to improve the multilingual instruction-following ability and downstream performance of large language models (LLMs), but the resulting multilingual capability remains highly sensitive to the composition and selection of the training data. Existing selection methods, often based on features like text quality, diversity, or task relevance, typically overlook the intrinsic linguistic structure of multilingual data. In this paper, we propose LangGPS, a lightweight two-stage pre-selection framework guided by language separability which quantifies how well samples in different languages can be distinguished in the model's representation space. LangGPS first filters training data based on separability scores and then refines the subset using existing selection methods. Extensive experiments across six benchmarks and 22 languages demonstrate that applying LangGPS on top of existing selection methods improves their effectiveness and generalizability in multilingual training, especially for understanding tasks and low-resource languages. Further analysis reveals that highly separable samples facilitate the formation of clearer language boundaries and support faster adaptation, while low-separability samples tend to function as bridges for cross-lingual alignment. Besides, we also find that language separability can serve as an effective signal for multilingual curriculum learning, where interleaving samples with diverse separability levels yields stable and generalizable gains. Together, we hope our work offers a new perspective on data utility in multilingual contexts and support the development of more linguistically informed LLMs.

多语言模型数据筛选语言可分性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。