arXiv:2506.04645cs.LGcs.DC2025-06被引 18

量化大模型推理的性价比,优化成本与速度的权衡。

Inference economics of language models

  • 构建理论模型分析成本与生成速度的经济权衡
  • 给出主流大模型在不同设置下的性价比最优解
  • 适合关注部署效率的工程师和研究者参考

我们提出一个理论模型,解决大规模部署语言模型时,每字节成本与串行生成速度之间的经济权衡问题。模型综合考虑算力、内存带宽、网络带宽及延迟约束,通过优化不同并行策略和批处理大小,找到在特定每字节成本下串行推理速度最优的配置。利用该模型,计算了多个主流语言模型的串行速度与每字节成本之间的帕累托前沿。

原文摘要 · Abstract (English)

We develop a theoretical model that addresses the economic trade-off between cost per token versus serial token generation speed when deploying LLMs for inference at scale. Our model takes into account arithmetic, memory bandwidth, network bandwidth and latency constraints; and optimizes over different parallelism setups and batch sizes to find the ones that optimize serial inference speed at a given cost per token. We use the model to compute Pareto frontiers of serial speed versus cost per token for popular language models.

大模型推理成本优化性能评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。