从系统视角重新审视推理时扩展,发现算力最优不等于实际最优。
Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- 提出系统驱动的测试时扩展评估框架,关注延迟与每字成本。
- 实测显示张量并行与推测解码优化存在实际瓶颈。
- 适合关注大模型部署效率的研究者与工程团队。
测试时扩展(TTS)近期被视为挖掘预训练大语言模型隐藏推理能力的有前景方向。然而,现有方法仅聚焦于计算最优的帕累托前沿,忽视了计算最优未必是系统最优这一事实。本文从系统驱动视角分析推理模型在延迟和每字成本等实际指标下的扩展特性。通过评估张量并行与推测解码等常见优化策略的影响,初步分析揭示了当前方法的局限性,呼吁向整体化、系统感知的评估范式转变,以真正捕捉推理时扩展规律的本质。
原文摘要 · Abstract (English)
Test-time scaling (TTS) has recently emerged as a promising direction to exploit the hidden reasoning capabilities of pre-trained large language models (LLMs). However, existing scaling methods narrowly focus on the compute-optimal Pareto-frontier, ignoring the simple fact that compute-optimal is not always system-optimal. In this work, we propose a system-driven perspective on TTS, analyzing how reasoning models scale against practical metrics, such as latency and cost-per-token. By evaluating the impact of popular optimizations such as tensor parallelism and speculative decoding, our preliminary analysis reveals the limitations of current methods and calls for a paradigm shift toward holistic, system-aware evaluations that capture the true essence of scaling laws at inference time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。