评估大模型无思考过程推理速度,预警安全监控风险
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

- 通过任务完成时间与推理令牌数,量化模型无CoT推理能力
- 过去六年模型无思考推理速度每年翻倍,GPT-5.5已超3分钟
- 建议开发者追踪此指标,尤其关注2028年后可能超7分钟的阈值
许多保障前沿AI模型安全的措施依赖于监控其链式思维(CoT)推理过程。若模型能在不显式输出思考过程的情况下完成复杂推理,将削弱此类监督机制。本文在43个领域、超过3万道题目上评估了前沿模型在无CoT情况下的推理表现,涵盖数学、编程、谜题、因果推理、心智理论及策略推理等。为对比人类表现,我们定义了50%任务完成时间阈值(TH):即模型在50%成功率下所需的人类完成时间。同时引入50%推理令牌阈值:模型以50%成功率完成任务所需的最少o3-mini推理令牌数。结果显示,过去六年中,前沿模型的无CoT TH约每年翻倍,当前GPT-5.5的TH已超3分钟,推理令牌阈值超过1,500。中位预测显示,到2028年无CoT TH可能突破7分钟,2030年可达25分钟,但存在较大不确定性。建议前沿开发者主动追踪该指标。
原文摘要 · Abstract (English)
Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficiently complex reasoning internally, without explicit thinking tokens, this would undermine such oversight. We measure how well frontier models reason without CoT across a suite of over 30,000 questions spanning 43 benchmarks in domains including math, coding, puzzles, causality, theory-of-mind, and strategic reasoning. To compare models against humans, we estimate the $50\%$-task-completion time horizon (TH): the human time required for tasks a model completes with $50\%$ success rate. We complement this with a $50\%$ reasoning token horizon: the minimum number of o3-mini reasoning tokens needed for tasks a model solves with $50\%$ success rate. We find that the no-CoT $50\%$ TH of frontier models has been doubling roughly every year over the past six years, with GPT-5.5's TH reaching over 3 minutes and reasoning token horizon exceeding 1,500 tokens. Our median estimates predict that frontier no-CoT THs could exceed 7 minutes by 2028, and 25 minutes by 2030, though these projections carry substantial uncertainty. We recommend frontier developers track this explicitly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。