首个针对数据流加速器的LLM基准测试框架,助力高效硬件优化。
DABench-LLM: Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators for LLMs
- 构建融合片内分析与片间扩展的综合评估方法
- 在3款主流数据流芯片上揭示性能瓶颈并提出优化策略
- 适合从事AI硬件设计与系统优化的研究者参考
大语言模型的指数级增长已超出传统CPU和GPU架构的承载能力,因摩尔定律放缓。数据流AI加速器成为有前景的替代方案,但缺乏对LLM训练的深入性能分析与标准化基准测试方法。本文提出DABench-LLM,首个专为数据流加速器上的LLM工作负载设计的基准测试框架。通过结合片内性能剖析与片间可扩展性分析,该框架可全面评估资源分配、负载均衡和资源效率等关键指标。它帮助研究人员快速洞察底层软硬件行为,并提供性能优化指导。我们在Cerebras WSE-2、SambaNova RDU和Graphcore IPU三款商用数据流加速器上验证了该框架,揭示了性能瓶颈并提出具体优化建议,证明其在多种数据流AI硬件平台上的通用性与有效性。
原文摘要 · Abstract (English)
The exponential growth of large language models has outpaced the capabilities of traditional CPU and GPU architectures due to the slowdown of Moore's Law. Dataflow AI accelerators present a promising alternative; however, there remains a lack of in-depth performance analysis and standardized benchmarking methodologies for LLM training. We introduce DABench-LLM, the first benchmarking framework designed for evaluating LLM workloads on dataflow-based accelerators. By combining intra-chip performance profiling and inter-chip scalability analysis, DABench-LLM enables comprehensive evaluation across key metrics such as resource allocation, load balance, and resource efficiency. The framework helps researchers rapidly gain insights into underlying hardware and system behaviors, and provides guidance for performance optimizations. We validate DABench-LLM on three commodity dataflow accelerators, Cerebras WSE-2, SambaNova RDU, and Graphcore IPU. Our framework reveals performance bottlenecks and provides specific optimization strategies, demonstrating its generality and effectiveness across a diverse range of dataflow-based AI hardware platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。