arXiv:2503.00461cs.ARcs.AI2025-03中稿 · appear at DATE 202…被引 4

用存内计算提升TPU生成模型推理效率,省电又提速。

Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs

  • 在TPU中用存内计算替代传统矩阵单元,提升能效。
  • 大语言模型和扩散模型推理速度最高提升44.2%。
  • 适合关注AI芯片能效与生成模型部署的研究者。

随着生成模型的快速发展,如何在专用硬件上高效部署成为关键问题。张量处理单元(TPUs)虽能加速人工智能任务,但高功耗限制了其应用。存内计算(CIM)因其卓越的面积与能效优势受到关注。本文提出一种集成数字存内计算的TPU架构,取代传统数字阵列的矩阵乘法单元(MXUs)。通过构建基于CIM的TPU架构模型与仿真器,评估其对多种生成模型推理的优化效果。基于设计洞察,进一步探索不同架构选择。相比基线TPUv4i架构,大语言模型与扩散变换器推理分别实现最高44.2%和33.8%的性能提升,且MXU能耗降低27.3倍。

原文摘要 · Abstract (English)

With the rapid advent of generative models, efficiently deploying these models on specialized hardware has become critical. Tensor Processing Units (TPUs) are designed to accelerate AI workloads, but their high power consumption necessitates innovations for improving efficiency. Compute-in-memory (CIM) has emerged as a promising paradigm with superior area and energy efficiency. In this work, we present a TPU architecture that integrates digital CIM to replace conventional digital systolic arrays in matrix multiply units (MXUs). We first establish a CIM-based TPU architecture model and simulator to evaluate the benefits of CIM for diverse generative model inference. Building upon the observed design insights, we further explore various CIM-based TPU architectural design choices. Up to 44.2% and 33.8% performance improvement for large language model and diffusion transformer inference, and 27.3x reduction in MXU energy consumption can be achieved with different design choices, compared to the baseline TPUv4i architecture.

TPU存内计算生成模型能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。