用合成思维过程训练大模型,跨领域推理能力显著提升。
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
- 通过重建文本背后的思维过程生成合成数据,用于持续预训练。
- 在MMLU上跨领域推理准确率提升,难题上最高增8分。
- 模型能根据题目难易自动调整推理深度,适合多领域应用。
大语言模型通过监督微调和强化学习在数学、编程等特定领域展现出强大推理能力,但这些方法受限于任务特定信号,难以扩展。相比之下,持续预训练(CPT)无需任务标注,但如何有效合成推理数据仍不明确。本文提出基于隐藏思维过程的推理型持续预训练(Reasoning CPT),利用来自科学与法律领域的文本重构作者思考路径,生成合成训练数据。以Gemma2-9B为实验模型,在MMLU基准上对比标准CPT,结果表明:推理型CPT在所有评估领域均表现更优;不同领域间的推理能力可迁移,且随着问题难度上升,性能差距扩大,最困难问题上最高提升8分;此外,模型能依据问题复杂度动态调整推理深度。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated significant improvements in reasoning capabilities through supervised fine-tuning and reinforcement learning. However, when training reasoning models, these approaches are primarily applicable to specific domains such as mathematics and programming, which imposes fundamental constraints on the breadth and scalability of training data. In contrast, continual pretraining (CPT) offers the advantage of not requiring task-specific signals. Nevertheless, how to effectively synthesize training data for reasoning and how such data affect a wide range of domains remain largely unexplored. This study provides a detailed evaluation of Reasoning CPT, a form of CPT that uses synthetic data to reconstruct the hidden thought processes underlying texts, based on the premise that texts are the result of the author's thinking process. Specifically, we apply Reasoning CPT to Gemma2-9B using synthetic data with hidden thoughts derived from STEM and Law corpora, and compare it to standard CPT on the MMLU benchmark. Our analysis reveals that Reasoning CPT consistently improves performance across all evaluated domains. Notably, reasoning skills acquired in one domain transfer effectively to others; the performance gap with conventional methods widens as problem difficulty increases, with gains of up to 8 points on the most challenging problems. Furthermore, models trained with hidden thoughts learn to adjust the depth of their reasoning according to problem difficulty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。