提出新方法优化多语言训练数据比例,提升大模型跨语言能力。
Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining
- 基于跨语言交互感知的比率计算,量化各语言有效贡献
- 两阶段优化策略使多语言性能显著提升,超越基线
- 适合需要高效多语言训练的AI研究者与工程师
大语言模型在全球应用中日益重要,其多语言能力的关键在于训练语料中语言比例的合理分配。然而,由于跨语言交互复杂且对数据规模敏感,最优比例难以确定。本文提出Climb(跨语言交互感知的多语言平衡框架),通过显式捕捉语言间依赖关系,量化每种语言的有效分配比例。Climb采用两阶段优化:首先均衡各语言的边际收益,再最大化分配向量的幅度,显著简化了复杂的多语言优化问题。大量实验表明,Climb能准确测量不同多语言场景下的跨语言交互;使用其推导比例训练的LLM在多语言性能上持续达到领先水平,甚至在总训练量较少的情况下,表现可媲美开源大模型。
原文摘要 · Abstract (English)
Large language models (LLMs) have become integral to a wide range of applications worldwide, driving an unprecedented global demand for effective multilingual capabilities. Central to achieving robust multilingual performance is the strategic allocation of language proportions within training corpora. However, determining optimal language ratios is highly challenging due to intricate cross-lingual interactions and sensitivity to dataset scale. This paper introduces Climb (Cross-Lingual Interaction-aware Multilingual Balancing), a novel framework designed to systematically optimize multilingual data allocation. At its core, Climb introduces a cross-lingual interaction-aware language ratio, explicitly quantifying each language's effective allocation by capturing inter-language dependencies. Leveraging this ratio, Climb proposes a principled two-step optimization procedure--first equalizing marginal benefits across languages, then maximizing the magnitude of the resulting language allocation vectors--significantly simplifying the inherently complex multilingual optimization problem. Extensive experiments confirm that Climb can accurately measure cross-lingual interactions across various multilingual settings. LLMs trained with Climb-derived proportions consistently achieve state-of-the-art multilingual performance, even achieving competitive performance with open-sourced LLMs trained with more tokens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。