arXiv:2505.06085cs.PFcs.AI2025-05中稿 · the Computational …被引 2

评测了Tenstorrent RISC-V加速器在低精度下的矩阵乘法性能,发现其能效比表现优异。

Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities

  • 使用RISC-V架构的Grayskull芯片,针对低精度矩阵乘法进行性能优化
  • 在BF16精度下达到1.55 TFLOPs/W的峰值能效,显著优于通用处理器
  • 适合关注能效比的边缘或嵌入式AI部署场景

生成式AI服务对大语言模型(LLMs)的需求激增,推动了专用硬件架构的发展以提升计算效率并降低能耗。本文评估了Tenstorrent Grayskull e75 RISC-V加速器在低数值精度下执行基础线性代数核函数(如矩阵乘法)的性能,该操作是LLM计算的核心。我们详细分析了执行模型、网格尺寸、矩阵维度、数据格式及数值精度对计算效率的影响。进一步将Grayskull与业界领先架构对比,包括Intel Sapphire Rapids处理器以及NVIDIA V100和A100 GPU。尽管NVIDIA GPU在原始性能上占优,但Grayskull在功耗与计算吞吐量之间展现出有竞争力的平衡,在BF16精度下实现了1.55 TFLOPs/W的峰值能效。

原文摘要 · Abstract (English)

The increasing demand for generative AI as Large Language Models (LLMs) services has driven the need for specialized hardware architectures that optimize computational efficiency and energy consumption. This paper evaluates the performance of the Tenstorrent Grayskull e75 RISC-V accelerator for basic linear algebra kernels at reduced numerical precision, a fundamental operation in LLM computations. We present a detailed characterization of Grayskull's execution model, gridsize, matrix dimensions, data formats, and numerical precision impact computational efficiency. Furthermore, we compare Grayskull's performance against state-of-the-art architectures with tensor acceleration, including Intel Sapphire Rapids processors and two NVIDIA GPUs (V100 and A100). Whilst NVIDIA GPUs dominate raw performance, Grayskull demonstrates a competitive trade-off between power consumption and computational throughput, reaching a peak of 1.55 TFLOPs/Watt with BF16.

RISC-V矩阵乘法能效比AI加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。