arXiv:2503.18866cs.LGcs.AI2025-03被引 42

通过挖掘文本背后的潜在思维过程,提升大模型在数据有限时的预训练效率。

Reasoning to Learn from Latent Thoughts

  • 将网页文本视为人类思维过程的压缩结果,利用潜在思维提升数据利用率。
  • 10亿参数模型通过自迭代优化,在数学任务上显著超越仅用原始数据训练的基线。
  • 无需强教师模型,可通过自举机制持续改进自身性能,适合资源受限场景。

语言模型预训练的计算规模已超过人类撰写文本的增长速度,引发数据成为瓶颈的担忧。为在数据受限条件下继续扩展预训练,我们提出显式建模和推断文本生成背后的潜在思维,可显著提升数据效率。直观上,网络文本是人类冗长思维过程的压缩结果,其中蕴含的重要上下文知识与推理步骤对高效学习至关重要。我们在数学任务的受限数据持续预训练中验证了该方法的有效性:合成数据推断潜在思维相比原始数据训练大幅提升数据效率;进一步展示无强教师指导下的自举能力——语言模型通过EM算法迭代优化自身性能与思维增强数据质量。10亿参数模型至少完成三轮自迭代,且随着推断计算增加,性能增益持续扩大,显著优于仅使用原始数据训练的基线模型。推断计算扩展与迭代次数带来的收益,揭示了数据受限预训练的新扩展路径。

原文摘要 · Abstract (English)

Compute scaling for language model (LM) pretraining has outpaced the growth of human-written texts, leading to concerns that data will become the bottleneck to LM scaling. To continue scaling pretraining in this data-constrained regime, we propose that explicitly modeling and inferring the \emph{latent thoughts} that underlie the text generation process can significantly improve pretraining data efficiency. Intuitively, our approach views web text as the compressed final outcome of a verbose human thought process and that the latent thoughts contain important contextual knowledge and reasoning steps that are critical to data-efficient learning. We empirically demonstrate the effectiveness of our approach through data-constrained continued pretraining for math. We first show that synthetic data approaches to inferring latent thoughts significantly improve data efficiency over training on the same amount of raw data. Furthermore, we demonstrate latent thought inference without a strong teacher, where an LM \emph{bootstraps its own performance} by using an EM algorithm to iteratively improve the capability of the trained LM and the quality of thought-augmented pretraining data. We show that a 1B LM can bootstrap its performance across at least three iterations and significantly outperform baselines trained on raw data, with increasing gains from additional inference compute when performing the E-step. The gains from inference scaling and EM iterations suggest new opportunities for scaling data-constrained pretraining.

大模型思维建模数据效率自举学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。