arXiv:2503.00322cs.ARcs.AI2025-03中稿 · IEEE ISSCC 2025

降低大模型推理内存访问,提升芯片效率

T-REX: A 68-567 μs/token, 0.41-3.95 μJ/token Transformer Accelerator with Reduced External Memory Access and Enhanced Hardware Utilization in 16nm FinFET

  • 动态批处理+双向寄存器文件减少内存读写
  • 每字节能耗0.41-3.95μJ,延迟68-567μs/token
  • 适合部署在高能效边缘设备的Transformer加速

本文提出新型训练与后训练压缩方案,以减少Transformer模型推理时的外部内存访问。同时引入一种名为动态批处理的新控制流机制,以及一种称为双向可访问寄存器文件的新型缓冲架构,进一步降低外部内存访问频率,提升硬件利用率。该设计在16nm FinFET工艺下实现68-567 μs/token的延迟,每字节能耗为0.41-3.95 μJ/token,显著优化了内存访问开销与芯片资源利用。

原文摘要 · Abstract (English)

This work introduces novel training and post-training compression schemes to reduce external memory access during transformer model inference. Additionally, a new control flow mechanism, called dynamic batching, and a novel buffer architecture, termed a two-direction accessible register file, further reduce external memory access while improving hardware utilization.

Transformer加速低功耗内存优化芯片设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。