arXiv:2505.14733cs.LGcs.AI2025-05被引 17

用推理时计算提升大模型能效,复杂推理更省电

The Energy Cost of Reasoning: Analyzing Energy Usage in LLMs with Test-time Compute

  • 推理时动态加算力,替代盲目扩大模型
  • 复杂推理任务下准确率提升且能耗更低
  • 按问题难易调整算力,适合部署优化

大语言模型的规模扩张虽推动显著进展,但面临收益递减和能源消耗激增的问题。本文探索测试时计算(Test-time Compute, TTC)作为传统扩容策略的节能补充:在推理阶段而非训练阶段分配额外计算资源。实验表明,相较于单纯增大模型规模,TTC在准确率与能耗之间实现更优平衡,尤其在需要复杂推理的任务中表现更优。此外,研究发现TTC性能与输出序列长度存在关键交互关系,通过根据查询复杂度动态调节推理阶段的计算资源,可显著提升效率。结果支持将TTC作为未来语言模型可持续、高精度、可适应部署的重要方向。

原文摘要 · Abstract (English)

Scaling large language models (LLMs) has driven significant advancements, yet it faces diminishing returns and escalating energy demands. This work explores how test-time compute (TTC) can serve as an energy-efficient complement to conventional scaling strategies by allocating additional computational resources at inference time rather than during training. Specifically, we investigate whether employing TTC can achieve superior accuracy-energy trade-offs compared to simply increasing model size. Our empirical analysis reveals that TTC surpasses traditional model scaling in accuracy/energy efficiency, with notable gains in tasks demanding complex reasoning rather than mere factual recall. Further, we identify a critical interaction between TTC performance and output sequence length, demonstrating that strategically adjusting compute resources at inference time according to query complexity can substantially enhance efficiency. Our findings advocate for TTC as a promising direction, enabling more sustainable, accurate, and adaptable deployment of future language models.

大模型能效优化推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。