用实测数据帮用户选最省钱高效的LLM推理硬件
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
- 在真实负载下测试不同GPU的LLM推理表现
- 比现有方法多33%满足性能要求,平均降本60%
- 适合部署LLM服务的工程师和决策者
随着大语言模型(LLMs)日益普及,其推理服务需在满足性能要求的前提下支持数千用户请求。服务性能主要取决于部署的硬件,但选择合适硬件以达成性能目标仍具挑战。本文提出首个此类系统 LLM-Pilot,通过在多种GPU上于真实工作负载下对LLM推理服务进行基准测试,并针对每种GPU优化服务配置以最大化性能。基于该表征数据,LLM-Pilot构建预测模型,可为未见过的LLM推荐最具成本效益的硬件。相比现有方法,该系统在性能达标率上提升33%,平均成本降低60%。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are rapidly growing in popularity, LLM inference services must be able to serve requests from thousands of users while satisfying performance requirements. The performance of an LLM inference service is largely determined by the hardware onto which it is deployed, but understanding of which hardware will deliver on performance requirements remains challenging. In this work we present LLM-Pilot - a first-of-its-kind system for characterizing and predicting performance of LLM inference services. LLM-Pilot performs benchmarking of LLM inference services, under a realistic workload, across a variety of GPUs, and optimizes the service configuration for each considered GPU to maximize performance. Finally, using this characterization data, LLM-Pilot learns a predictive model, which can be used to recommend the most cost-effective hardware for a previously unseen LLM. Compared to existing methods, LLM-Pilot can deliver on performance requirements 33% more frequently, whilst reducing costs by 60% on average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。