评测大模型在多种硬件上的推理性能,帮用户选最优配置。
LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators
- 构建跨平台基准测试工具,覆盖主流GPU和AI加速器。
- 对比7B与70B参数模型在不同硬件上的吞吐表现差异。
- 提供交互式仪表板,辅助快速定位最佳推理配置。
大型语言模型(LLMs)在多个领域推动了突破性进展,广泛应用于文本生成。然而,这些复杂模型的计算需求带来了巨大挑战,亟需高效的硬件加速。对LLM在多样化硬件平台上的性能进行基准测试,对于理解其可扩展性和吞吐特性至关重要。本文提出LLM-Inference-Bench,一个全面的基准测试套件,用于评估LLM在硬件上的推理性能。我们系统分析了包括Nvidia和AMD GPU,以及Intel Habana和SambaNova等专用AI加速器在内的多种硬件平台。评估涵盖来自LLaMA、Mistral和Qwen系列的7B和70B参数模型及多个LLM推理框架。基准测试结果揭示了不同模型、硬件平台和推理框架的优势与局限。我们还提供了交互式仪表板,帮助用户针对特定硬件平台识别最优性能配置。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have propelled groundbreaking advancements across several domains and are commonly used for text generation applications. However, the computational demands of these complex models pose significant challenges, requiring efficient hardware acceleration. Benchmarking the performance of LLMs across diverse hardware platforms is crucial to understanding their scalability and throughput characteristics. We introduce LLM-Inference-Bench, a comprehensive benchmarking suite to evaluate the hardware inference performance of LLMs. We thoroughly analyze diverse hardware platforms, including GPUs from Nvidia and AMD and specialized AI accelerators, Intel Habana and SambaNova. Our evaluation includes several LLM inference frameworks and models from LLaMA, Mistral, and Qwen families with 7B and 70B parameters. Our benchmarking results reveal the strengths and limitations of various models, hardware platforms, and inference frameworks. We provide an interactive dashboard to help identify configurations for optimal performance for a given hardware platform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。