用可控语言实验揭示模型跨语言迁移的关键机制。
An In-Vitro Study on Cross-Lingual Generalization in Language Models
- 构建两种表面不同但结构相同的生成语言,分离变量研究迁移。
- 小词表更利于迁移,因能保留可复用的语义片段。
- 语法和类型能力先于词汇迁移出现,适合研究语言本质。
语言模型的跨语言迁移在真实语料中难以研究,因词汇重叠、形态、数据不平衡与分词混杂。本文提出一种体外实验框架,使用两种具有相同本体、类型语法和组合结构但表面实现不同的程序生成语言。可独立调控词汇距离、少数语言比例、分词器训练方式与词表大小,评估在未见过词汇形式的掩码少数语言条件下的迁移表现。700次受控实验表明,迁移效果主要取决于分词是否保留可复用的跨语言子结构,而非分词平衡或原始词汇相似性。小词表常提升掩码迁移性能,因能将词语分解为共享片段;大词表则使形式成为语言特有原子。此外,迁移呈现阶段性:语法与类型能力先于掩码词汇泛化出现。我们进一步通过分词桥机制解释该现象,发现桥接强度与掩码可达性显著相关。
原文摘要 · Abstract (English)
Cross-lingual transfer in language models is difficult to study in natural corpora because lexical overlap, morphology, data imbalance, and tokenization are entangled. We introduce an in-vitro framework with two procedurally generated languages that share the same ontology, typed grammar, and compositional structure, but differ in surface realization. This lets us independently vary lexical distance, minority-language proportion, tokenizer training regime, and vocabulary size, while evaluating transfer on a masked minority-language condition whose lexical forms are never observed during training. Across 700 controlled runs, we find that transfer is governed less by tokenizer balance or raw lexical similarity than by whether tokenization preserves reusable cross-lingual substructure. Smaller vocabularies often improve masked transfer by keeping words decomposable into shared fragments, whereas larger vocabularies can turn forms into language-specific atoms. We further show that transfer emerges as a staged process: grammatical and type-level competence precede masked lexical generalization. Finally, we attempt to explain this mechanism through tokenizer bridges and show that bridge strength correlates strongly with masked reachability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。