对比高通与英伟达芯片,发现高通在低功耗下可高效运行大模型。
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
- 用vLLM框架测试12个大模型,比较芯片能效与扩展性。
- 700亿参数模型仅需1张高通卡,功耗比8张英伟达卡低20倍。
- 小模型单卡功耗仅为英伟达4卡配置的1/35,适合节能场景。
本研究对高通Cloud AI 100 Ultra加速器在大型语言模型(LLM)推理中的表现进行了基准测试,评估其在国家研究平台(NRP)生态系统中与英伟达A100 GPU(4x和8x配置)的能效、性能及硬件可扩展性。共部署12个开源LLM,参数量从1.24亿到700亿不等,均使用vLLM框架。结果显示,高通芯片在特定模型上实现有竞争力的能效,支持更细粒度的硬件分配:部分700亿参数模型仅需1张高通卡即可运行,而英伟达需8张,功耗从2,983W降至148W,降低20倍;对于小型模型,单张高通卡功耗仅36W,相较4张英伟达卡的1,246W降低35倍。研究为能源受限的高性能计算部署提供了新选择。
原文摘要 · Abstract (English)
This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throughput per watt), performance, and hardware scalability against NVIDIA A100 GPUs (in 4x and 8x configurations) within the National Research Platform (NRP) ecosystem. A total of 12 open-source LLMs, ranging from 124 million to 70 billion parameters, are served using the vLLM framework. Our analysis reveals that QAic achieves competitive energy efficiency with advantages on specific models while enabling more granular hardware allocation: some 70B models operate on as few as 1 QAic card versus 8 A100 GPUs required, with 20x lower power consumption (148W vs 2,983W). For smaller models, single QAic devices achieve up to 35x lower power consumption compared to our 4-GPU A100 configuration (36W vs 1,246W). The findings offer insights into the potential of the Qualcomm Cloud AI 100 Ultra for energy-constrained and resource-efficient HPC deployments within the National Research Platform (NRP).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。