让大模型推理时可快可慢,灵活切换思考模式。
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
- 用一个参数控制推理过程,动态调节慢思考转快思考。
- 在数学、编程、科学任务上表现优于现有方法。
- 适合需要灵活推理速度的AI应用开发人员。
本文提出AlphaOne(α1),一种在测试阶段调节大型推理模型(LRMs)推理进程的通用框架。α1引入α时刻,以统一参数α表示缩放后的思考阶段。在该缩放前α时刻阶段,通过将推理转换令牌插入建模为伯努利随机过程,动态调度慢思考过渡。α时刻后,α1通过结束思考令牌确定性终止慢思考,从而促进快速推理和高效答案生成。该方法统一并泛化了现有单调缩放方法,实现了灵活且密集的慢速到快速推理调节。在数学、编程和科学等多个挑战性基准上的广泛实证研究显示,α1具备卓越的推理能力和效率。
原文摘要 · Abstract (English)
This paper presents AlphaOne ($α$1), a universal framework for modulating reasoning progress in large reasoning models (LRMs) at test time. $α$1 first introduces $α$ moment, which represents the scaled thinking phase with a universal parameter $α$. Within this scaled pre-$α$ moment phase, it dynamically schedules slow thinking transitions by modeling the insertion of reasoning transition tokens as a Bernoulli stochastic process. After the $α$ moment, $α$1 deterministically terminates slow thinking with the end-of-thinking token, thereby fostering fast reasoning and efficient answer generation. This approach unifies and generalizes existing monotonic scaling methods by enabling flexible and dense slow-to-fast reasoning modulation. Extensive empirical studies on various challenging benchmarks across mathematical, coding, and scientific domains demonstrate $α$1's superior reasoning capability and efficiency. Project page: https://alphaone-project.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。