用多领域数据训练大模型,让AI推理更准更省
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
- 融合科学、人文等多领域问答数据,构建跨任务推理能力
- 数学题准确率提升30.1%,非数学题最高增15.1%
- 适合想提升模型泛化与推理效率的研究者和开发者
大型语言模型在强化学习(RL)增强下展现出强大推理能力,尤其在数学推理领域表现突出。然而,将此类方法推广到更广泛的推理任务仍面临数据有限、可验证奖励机制缺失及任务多样性等挑战。本文提出NEMOTRON-CROSSTHINK框架,通过系统整合多领域语料(包括合成与真实问答对),提升模型在多样化推理任务中的泛化能力。该框架通过四方面优化:(1) 融合涵盖STEM、人文学科、社会科学等多源数据;(2) 使用选择题与开放题结构化模板控制答案空间复杂度;(3) 过滤可验证答案;(4) 优化多源数据混合策略。实验表明,该方法在数学基准上显著提升性能(MATH-500: +30.1%,AMC23: +27.5%),非数学任务亦有显著进步(MMLU-PRO: +12.8%,GPQA-DIAMOND: +11.3%,AGIEVAL: +15.1%,SUPERGPQA: +3.8%)。同时,模型响应效率提升,正确回答平均减少28%的令牌消耗,体现更聚焦高效的推理过程。结果证明,基于多领域、多格式数据的强化学习训练,可实现更准确、高效且通用的大型语言模型。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown strong reasoning capabilities, particularly when enhanced through Reinforcement Learning (RL). While prior work has successfully applied RL to mathematical reasoning -- where rules and correctness are well-defined -- generalizing these methods to broader reasoning domains remains challenging due to limited data, the lack of verifiable reward structures, and diverse task requirements. In this work, we propose NEMOTRON-CROSSTHINK, a framework that systematically incorporates multi-domain corpora, including both synthetic and real-world question-answer pairs, into RL training to improve generalization across diverse reasoning tasks. NEMOTRON-CROSSTHINK addresses key challenges by (1) incorporating data from varied sources spanning STEM, humanities, social sciences, etc.; (2) applying structured templates (e.g., multiple-choice and open-ended) to control answer-space complexity; (3) filtering for verifiable answers; and (4) optimizing data blending strategies that utilizes data from multiple sources effectively. Our approach enables scalable and verifiable reward modeling beyond mathematics and demonstrates improved accuracies on both math (MATH-500: +30.1%, AMC23:+27.5%) and non-math reasoning benchmarks (MMLU-PRO: +12.8%, GPQA-DIAMOND: +11.3%, AGIEVAL: +15.1%, SUPERGPQA: +3.8%). Moreover, NEMOTRON-CROSSTHINK exhibits significantly improved response efficiency -- using 28% fewer tokens for correct answers -- highlighting more focused and effective reasoning. Through NEMOTRON-CROSSTHINK, we demonstrate that integrating multi-domain, multi-format data in RL leads to more accurate, efficient, and generalizable LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。