arXiv:2410.01720cs.AIcs.CL2024-10ICLR被引 24

从反瓶颈视角揭示合成数据如何影响大模型泛化能力

Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

  • 提出反瓶颈视角,分析生成模型的信息增益对泛化的影响
  • 定义通用性增益(GGMI),量化合成数据与泛化的关系
  • 为合成数据生成和后训练优化提供理论指导,适合模型开发者

由于高质量特定数据稀缺,合成数据已成为大语言模型后训练的关键资源。尽管已有多种生成方法,但其实际效果与理论理解之间仍存在显著差距。本文首先详细建模主流合成数据生成流程,基于此提出全新反瓶颈视角,证明后训练模型的泛化能力关键取决于生成模型带来的信息增益。进一步引入广义信息增益(GGMI)概念,阐明泛化增益与信息增益的关系。该分析为合成数据生成提供了理论基础,并揭示其与模型泛化能力的内在联系,有助于优化合成数据生成策略与后训练过程。代码已开源。

原文摘要 · Abstract (English)

Synthetic data has become a pivotal resource in post-training tasks for large language models (LLMs) due to the scarcity of high-quality, specific data. While various methods have been developed to generate synthetic data, there remains a discernible gap between the practical effects of synthetic data and our theoretical comprehension. To address this challenge, we commence by presenting a detailed modeling of the prevalent synthetic data generation process. Building upon this modeling, we demonstrate that the generalization capability of the post-trained model is critically determined by the information gain derived from the generative model, as analyzed from a novel reverse-bottleneck perspective. Moreover, we introduce the concept of Generalization Gain via Mutual Information (GGMI) and elucidate the relationship between generalization gain and information gain. This analysis serves as a theoretical foundation for synthetic data generation and further highlights its connection with the generalization capability of post-trained models, offering an understanding about the design of synthetic data generation techniques and the optimization of the post-training process. We open-source our code at https://github.com/ZyGan1999/Towards-a-Theoretical-Understanding-of-Synthetic-Data-in-LLM-Post-Training.

大模型合成数据泛化能力理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。