对比Gaudi-2与NVIDIA A100,验证其在能效和性能上的竞争力
Debunking the CUDA Myth Towards GPU-based AI Systems
- 通过微基准测试对比Gaudi-2与A100的计算、内存与通信性能
- Gaudi-2在端到端AI任务中实现与A100相当的能效表现
- 适合关注国产加速器替代方案及系统集成的开发者
本文全面评估了英特尔Gaudi NPU作为当前主流NVIDIA GPU的替代方案。首先,我们构建了一套微基准测试,对比Gaudi-2与NVIDIA A100,结果表明Gaudi-2在基础AI计算、内存访问和通信操作上表现具有竞争力,并在多个重要AI工作负载上实现了端到端的可比性能。其次,我们评估了Gaudi NPU的可编程性,探讨了针对关键FBGEMM算子和vLLM的软件优化策略,其效率与GPU优化版本相当。结果显示,Gaudi-2在能效方面与A100接近,但软件生态成熟度仍有提升空间。总体而言,若能有效集成至高级别AI框架,Gaudi NPUs有望挑战NVIDIA GPU在AI服务器市场的主导地位,但仍需进一步改进以全面媲美其成熟的软件体系。
原文摘要 · Abstract (English)
This paper presents a comprehensive evaluation of Intel Gaudi NPUs as an alternative to NVIDIA GPUs, which is currently the de facto standard in AI system design. First, we create a suite of microbenchmarks to compare Intel Gaudi-2 with NVIDIA A100, showing that Gaudi-2 achieves competitive performance not only in primitive AI compute, memory, and communication operations but also in executing several important AI workloads end-to-end. We then assess Gaudi NPU's programmability by discussing several software-level optimization strategies to employ for implementing critical FBGEMM operators and vLLM, evaluating their efficiency against GPU-optimized counterparts. Results indicate that Gaudi-2 achieves energy efficiency comparable to A100, though there are notable areas for improvement in terms of software maturity. Overall, we conclude that, with effective integration into high-level AI frameworks, Gaudi NPUs could challenge NVIDIA GPU's dominance in the AI server market, though further improvements are necessary to fully compete with NVIDIA's robust software ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。