arXiv:2409.12994cs.ARcs.AI2024-09被引 6

对比主流加速器在大模型训练中的性能与功耗表现。

Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML

  • 构建CARAML基准套件,自动化评估多种硬件上模型训练效率。
  • 涵盖NVIDIA、AMD、Graphcore等多款加速器,支持可复现测试。
  • 适合芯片设计者和系统优化者参考,评估硬件真实效能。

机器学习技术的快速发展推动了专用硬件加速器的发展,以提升模型训练效率。本文提出CARAML基准套件,用于评估基于Transformer的大语言模型与计算机视觉模型在NVIDIA、AMD和Graphcore等硬件加速器上的训练性能与能耗表现。CARAML提供一个轻量、自动化、可扩展且可复现的框架,支持对新型硬件架构上ML工作负载的性能与能效进行系统性评测。文中详细介绍了CARAML的设计与实现,以及配套的自定义功耗测量工具jpwr。

原文摘要 · Abstract (English)

The rapid advancement of machine learning (ML) technologies has driven the development of specialized hardware accelerators designed to facilitate more efficient model training. This paper introduces the CARAML benchmark suite, which is employed to assess performance and energy consumption during the training of transformer-based large language models and computer vision models on a range of hardware accelerators, including systems from NVIDIA, AMD, and Graphcore. CARAML provides a compact, automated, extensible, and reproducible framework for assessing the performance and energy of ML workloads across various novel hardware architectures. The design and implementation of CARAML, along with a custom power measurement tool called jpwr, are discussed in detail.

性能评估硬件加速能效分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。