arXiv:2503.05951cs.ARcs.AI2025-03被引 7

用大模型自动设计专用张量加速器,大幅降低功耗和面积

TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator

  • 基于大模型与检索增强生成,自动生成张量处理器架构
  • 相比人工优化,面积减少92%,功耗降低96%
  • 适合芯片设计自动化、AI加速器研发人员使用

深度神经网络的复杂性和规模不断增长,亟需专用张量加速器(如TPU)以满足计算与能效需求。然而,设计最优TPU仍面临领域专业知识门槛高、手工设计耗时长、缺乏高质量领域数据集等挑战。本文提出TPU-Gen,首个基于大语言模型(LLM)的框架,用于自动化生成精确与近似张量处理单元,聚焦流水线阵列架构。该框架配备精心构建的开源数据集,涵盖多种空间阵列设计及近似乘累加单元,支持针对不同DNN工作负载的设计复用与定制。通过引入检索增强生成(RAG),有效缓解硬件数据稀缺场景下大模型的幻觉问题。TPU-Gen将高层架构规范转化为优化的低层实现,实验表明其在面积和功耗上平均分别较人工优化基准降低92%和96%,为下一代由大模型驱动的设计自动化工具树立新标杆。

原文摘要 · Abstract (English)

The increasing complexity and scale of Deep Neural Networks (DNNs) necessitate specialized tensor accelerators, such as Tensor Processing Units (TPUs), to meet various computational and energy efficiency requirements. Nevertheless, designing optimal TPU remains challenging due to the high domain expertise level, considerable manual design time, and lack of high-quality, domain-specific datasets. This paper introduces TPU-Gen, the first Large Language Model (LLM) based framework designed to automate the exact and approximate TPU generation process, focusing on systolic array architectures. TPU-Gen is supported with a meticulously curated, comprehensive, and open-source dataset that covers a wide range of spatial array designs and approximate multiply-and-accumulate units, enabling design reuse, adaptation, and customization for different DNN workloads. The proposed framework leverages Retrieval-Augmented Generation (RAG) as an effective solution for a data-scare hardware domain in building LLMs, addressing the most intriguing issue, hallucinations. TPU-Gen transforms high-level architectural specifications into optimized low-level implementations through an effective hardware generation pipeline. Our extensive experimental evaluations demonstrate superior performance, power, and area efficiency, with an average reduction in area and power of 92\% and 96\% from the manual optimization reference values. These results set new standards for driving advancements in next-generation design automation tools powered by LLMs.

芯片设计大模型应用硬件生成加速器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。