用强化学习提升推理能力,小模型实现顶尖表现
Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
- 采用强化学习训练思维链,先推理再作答
- 在AIME、Codeforces等任务上达86.7、55.0分
- 200亿参数中仅激活20亿,适合高效部署
我们提出Seed1.5-Thinking,通过思考后再回答的方式提升推理能力,在多个基准测试中表现优异。该模型在AIME 2024上取得86.7分,Codeforces上达55.0分,GPQA上为77.3分,展现出在科学、技术、工程与数学及编程领域的强大推理能力。此外,其在非推理类任务上也表现出色,相比DeepSeek R1在胜率上高出8%。作为一项轻量级的Mixture-of-Experts(MoE)模型,总参数量200B,但仅激活20B,效率更高。为推动通用推理研究,我们还构建了两个内部评估基准:BeyondAIME和Codeforces,将公开发布以支持后续研究。模型试用链接:https://www.volcengine.com/experience/ark。
原文摘要 · Abstract (English)
We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 86.7 on AIME 2024, 55.0 on Codeforces and 77.3 on GPQA, demonstrating excellent reasoning abilities in STEM and coding. Beyond reasoning tasks, the method demonstrates notable generalization across diverse domains. For instance, it surpasses DeepSeek R1 by 8% in win rate on non-reasoning tasks, indicating its broader applicability. Compared to other state-of-the-art reasoning models, Seed1.5-Thinking is a Mixture-of-Experts (MoE) model with a relatively small size, featuring 20B activated and 200B total parameters. As part of our effort to assess generalized reasoning, we develop two internal benchmarks, BeyondAIME and Codeforces, both of which will be publicly released to support future research. Model trial link: https://www.volcengine.com/experience/ark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。