对比多种AI芯片推理性能,为大模型部署选型提供量化依据。
AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
- 从内存层级到互连结构,系统分析多类加速器架构差异。
- 批处理大小和序列长度变化导致性能最高相差3.7倍。
- 专家并行虽效率高但延迟波动更大,适合不同场景需求。
大语言模型的快速扩张正推动专用推理硬件的发展。本文首次开展面向工作负载、跨架构的商用AI加速器性能研究,涵盖基于GPU的芯片、混合封装与晶圆级引擎。通过对比内存层次、计算结构与片上互连,发现随着批处理大小和序列长度变化,不同架构间性能差异可达3.7倍。文中还评估了四种适用于万亿参数模型的扩展策略:专家并行在参数-算力比上具有8.4倍优势,但相比张量并行带来2.1倍更高的延迟波动。研究结果为工作负载与硬件匹配提供量化指导,并揭示下一代设计需弥补的架构短板。
原文摘要 · Abstract (English)
The rapid growth of large-language models (LLMs) is driving a new wave of specialized hardware for inference. This paper presents the first workload-centric, cross-architectural performance study of commercial AI accelerators, spanning GPU-based chips, hybrid packages, and wafer-scale engines. We compare memory hierarchies, compute fabrics, and on-chip interconnects, and observe up to 3.7x performance variation across architectures as batch size and sequence length change. Four scaling techniques for trillion-parameter models are examined; expert parallelism offers an 8.4x parameter-to-compute advantage but incurs 2.1x higher latency variance than tensor parallelism. These findings provide quantitative guidance for matching workloads to accelerators and reveal architectural gaps that next-generation designs must address.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。