arXiv:2502.16721cs.CLcs.AI2025-02被引 10

分析开源大模型在GPU上的实际速度,发现任务类型比每秒词数更重要。

Speed and Conversational Large Language Models: Not All Is About Tokens per Second

  • 对比主流开源LLM在不同任务下的运行速度
  • 揭示任务类型对推理速度影响大于吞吐量指标
  • 适合关注模型部署效率的研究者和开发者

本文研究了在GPU上运行时,开源权重大语言模型(LLMs)的速度及其与任务类型的依赖关系,对最流行的开源LLM进行了速度的对比分析。结果表明,模型的实际运行速度不仅取决于计算吞吐量,还显著受具体任务特性影响,单一以每秒处理词数(tokens per second)衡量性能存在局限。该研究强调需结合任务上下文评估模型效率,为模型选择与部署提供更全面的参考依据。

原文摘要 · Abstract (English)

The speed of open-weights large language models (LLMs) and its dependency on the task at hand, when run on GPUs, is studied to present a comparative analysis of the speed of the most popular open LLMs.

大模型推理速度模型部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。