arXiv:2510.16096cs.CLcs.LG2025-10被引 2

通过可控实验发现,预训练多样性影响语言模型的泛化能力。

Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization

  • 设计合成数据集,独立调控上下文结构与多样性水平。
  • 高多样性延迟事实准确率,但对跨分布泛化有关键作用。
  • 嵌入与解码层是影响跨分布性能的关键瓶颈。

语言模型在混合统计规律与特定词元间事实关联的序列上进行预训练。尽管近期研究指出事实关联的可变性(如改写)对泛化能力至关重要,但缺乏系统分析。本文提出一个灵活的合成测试平台,结合通用词元的统计流与源-目标词元对的事实流,实现对交互方式的细粒度控制。通过调整流的组成(上下文结构)和事实出现在哪些统计流中,可独立控制多样性类型与水平。受控实验表明,更高上下文多样性虽延迟分布内(ID)事实准确性,但对分布外(OOD)事实泛化的提升取决于上下文结构:某些情况下表现趋势一致,另一些情况下多样性成为非平凡事实召回的必要条件。即使低多样性导致事实召回失败,最优多样性水平也随训练时长变化。此外,我们识别出统计泛化独立失效或两者同时退化的结构。通过组件级干预,发现OOD失败源于不同优化瓶颈,凸显嵌入与解码层的重要性。该合成框架可隔离大规模研究中难以区分的效应,为未来研究提供可控测试平台。

原文摘要 · Abstract (English)

Language models are pretrained on sequences that blend statistical regularities (making text fluent) with factual associations between specific tokens (knowledge of facts). While recent work suggests that the variability of their interaction, such as paraphrases of factual associations, critically determines generalization ability, we lack a systematic analysis of these impacts. This paper introduces a flexible synthetic testbed that combines a statistical stream of generic tokens with an abstract factual stream of source-target token pairs, enabling fine-grained control over their interaction. The design enables the independent control of diversity nature by manipulating stream composition (contextual structure) and the diversity level by varying which statistical streams each fact appears in. Through controlled experiments, we find that while higher contextual diversity delays in-distribution (ID) factual accuracy, its impact on out-of-distribution (OOD) factual generalization depends critically on contextual structure. In some cases, OOD performance follows the same trend as ID, but in others, diversity becomes essential for non-trivial factual recall. Even when low diversity prohibits factual recall, optimal diversity levels depend on training duration. Beyond factual recall failures, we identify structures where statistical generalization fails independently, and others where both capabilities degrade. This shows how the interplay between contextual design and diversity level impacts different generalization aspects. Further, through a series of controlled interventions on the model components, we trace the OOD failures to distinct optimization bottlenecks, highlighting the importance of the embedding and unembedding layers. Our synthetic framework allows us to isolate effects that would be confounded in large-scale studies, offering a controlled testbed for future investigations.

语言模型泛化能力预训练多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。