测试多种语言后发现,预训练大模型的伪语言方法并不总有效。
Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

- 在多种语言上用不同分词器和模型大小验证伪预训练效果
- 仅小模型配合特定分词器时有稳定33%的训练效率提升
- 强调应多次实验避免采用不稳定的训练方法
在多种语言上对大模型进行伪语言预训练(pre-pretraining)是否能提升训练效率,此前报道称可节省33%训练令牌。本文在涵盖四个语系的多语言数据集上,使用两种分词器和不同模型规模进行了验证。结果表明,该增益高度依赖实验设置与随机种子;仅有128-Dyck伪语言配合Llama分词器在小型模型上对多数语言表现出稳定收益。我们进一步将效率变化与语言特征(如句长、形态丰富度、依存句法树深度、交叉依赖数等)相关联。研究指出,为避免社区采纳不稳定方法,至少应对部分实验进行多次重复训练。
原文摘要 · Abstract (English)
Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33% of training tokens needed to reach a certain performance. We validate this prior result for English on a larger set of natural languages across four language families, using two different tokenizers and varying model sizes. We also relate the observed gains (or losses) in token efficiency to quantified linguistic properties of the languages, such as sentence length, morphological richness, and features of dependency syntactic trees (tree depth, number of children, number of crossing dependencies). Our empirical results indicate that the reported gains depend heavily on the experiment setup and the choice of random seed, although we can confirm the trend of stable gains with 128-Dyck pretraining of small models with the Llama tokenizer for most of the examined languages. On a general note, we argue that multiple training runs should be carried out at least for a subset of experiments to avoid the community adopting unstable approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。