通过100万条长思维链数据训练,红星模型显著提升复杂推理能力。
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
- 构建红星模型,用百万级长思维链数据训练慢思考系统。
- 在MATH-Hard上从66.2%提升至81.6%,AIME解题率达46.7%。
- 小模型仅用数千样本即见效,适合需要强推理的场景。
本文探索了扩大长思维链(Long-CoT)数据规模至100万样本的潜力,开创性地开发出慢思考模型RedStar。通过在多种LLM及不同规模下进行大量实验,我们揭示了长思维链训练的优化要素与扩展规律。令人意外的是,即使小模型在有限数据下也表现显著提升,表明长思维链具备高样本效率,且样本难度在学习中起关键作用。研究发现,仅需数千样本即可有效触发长思维链推理,而大模型则实现质的飞跃。我们还提出强化学习(RL)规模训练作为推进慢思考系统的新方向。RedStar在多领域表现优异:在MATH-Hard上,RedStar-code-math性能从66.2%提升至81.6%;在美国数学奥林匹克(AIME)上,使用仅21k混合代码-数学数据集便解决46.7%的问题。在多模态任务如GeoQA和MathVista-GEO中,RedStar-Geo以极少长思维链数据取得竞争力结果,优于QvQ-Preview等其他慢思考系统。相较QwQ,RedStar在推理能力与泛化性间取得更佳平衡。研究表明,经精细调优,长思维链规模化可激发惊人推理性能,即便数据量有限,也为慢思考模型树立新标准。数据与模型已开源:https://huggingface.co/RedStar-Reasoning。
原文摘要 · Abstract (English)
Can scaling transform reasoning? In this work, we explore the untapped potential of scaling Long Chain-of-Thought (Long-CoT) data to 1000k samples, pioneering the development of a slow-thinking model, RedStar. Through extensive experiments with various LLMs and different sizes, we uncover the ingredients for specialization and scale for Long-CoT training. Surprisingly, even smaller models show significant performance gains with limited data, revealing the sample efficiency of Long-CoT and the critical role of sample difficulty in the learning process. Our findings demonstrate that Long-CoT reasoning can be effectively triggered with just a few thousand examples, while larger models achieve unparalleled improvements. We also introduce reinforcement learning (RL)-scale training as a promising direction for advancing slow-thinking systems. RedStar shines across domains: on the MATH-Hard benchmark, RedStar-code-math boosts performance from 66.2\% to 81.6\%, and on the USA Math Olympiad (AIME), it solves 46.7\% of problems using only 21k mixed-code-math datasets. In multimodal tasks like GeoQA and MathVista-GEO, RedStar-Geo achieves competitive results with minimal Long-CoT data, outperforming other slow-thinking systems like QvQ-Preview. Compared to QwQ, RedStar strikes the perfect balance between reasoning and generalizability. Our work highlights that, with careful tuning, scaling Long-CoT can unlock extraordinary reasoning capabilities-even with limited dataset and set a new standard for slow-thinking models across diverse challenges. Our data and models are released at https://huggingface.co/RedStar-Reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。