揭示大模型推理时计算扩展的极限,提出资源分配优化方法。
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
- 从概率建模出发,分析并行与串行推理扩展的理论边界。
- 发现两种扩展方式在计算预算上均有饱和点,超限则收益递减。
- 在AIME、MATH-500等基准上验证理论,适用于高效推理设计。
大型推理模型(LRMs)展现出通过内部测试时扩展提升推理能力的潜力。本文进一步探索测试时计算扩展的极限,提出测试时扩展性能模型(TTSPM)。从概率建模视角,理论分析了并行与串行两种扩展范式,推导出两者在计算预算上的饱和点,识别出额外计算带来边际收益递减的临界阈值。令人惊讶的是,尽管机制不同,两种范式在上界处收敛为统一数学结构。我们在AIME、MATH-500和GPQA等高难度推理基准上实证验证了理论结果,证明该边界对测试时资源分配具有实际指导意义。本工作旨在揭示测试时扩展的成本效益权衡,推动更高效的大型推理模型推理策略发展。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) have exhibited the capacity of enhancing reasoning performance via internal test-time scaling. Building upon this, a promising direction is to further scale test-time compute to unlock even greater reasoning capabilities. However, as we push these scaling boundaries, systematically understanding the practical limits and achieving optimal resource allocation becomes a critical challenge. In this paper, we investigate the scaling plateau of test-time scaling and introduce the Test-Time Scaling Performance Model (TTSPM). We theoretically analyze two fundamental paradigms for such extended scaling, parallel scaling and sequential scaling, from a probabilistic modeling perspective. Our primary contribution is the derivation of the saturation point on the scaling budget for both strategies, identifying thresholds beyond which additional computation yields diminishing returns. Remarkably, despite their distinct mechanisms, both paradigms converge to a unified mathematical structure in their upper bounds. We empirically validate our theoretical findings on challenging reasoning benchmarks, including AIME, MATH-500, and GPQA, demonstrating the practical utility of these bounds for test-time resource allocation. We hope that this work provides insights into the cost-benefit trade-offs of test-time scaling, guiding the development of more resource-efficient inference strategies for large reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。