arXiv:2511.06313cs.ARcs.AI2025-11被引 2

提出混合精度缩放树,提升NPU中低精度计算能效与灵活性。

Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration

  • 设计混合精度缩放树,兼顾整数与浮点累加优势
  • 在MXINT8/FP8/FP4下实现657至4065 GOPS/W能效
  • 集成8×8阵列到SNAX平台,适合高效神经网络部署

新兴的持续学习应用需要下一代神经处理单元(NPU)平台同时支持训练与推理。备受关注的Microscaling(MX)标准可为推理提供窄位宽,为训练提供大动态范围。然而,现有MX乘累加(MAC)设计面临关键权衡:整数累加需昂贵的窄浮点积转换,而FP32累加则存在量化损失和高昂归一化开销。为此,我们提出一种混合精度可扩展缩减树,结合两种方法优势,实现高效的混合精度累加并可控地放宽精度要求。此外,我们将8×8个此类MAC集成到当前最先进的NPU集成平台SNAX中,以实现对优化后的精度可扩展MX数据通路的高效控制与数据传输。我们在MAC级与系统级评估该设计,并与现有最优方案对比。集成系统在MXINT8、MXFP8/6和MXFP4下的能效分别为657、1438–1675和4065 GOPS/W,吞吐量分别为64、256和512 GOPS。

原文摘要 · Abstract (English)

Emerging continual learning applications necessitate next-generation neural processing unit (NPU) platforms to support both training and inference operations. The promising Microscaling (MX) standard enables narrow bit-widths for inference and large dynamic ranges for training. However, existing MX multiply-accumulate (MAC) designs face a critical trade-off: integer accumulation requires expensive conversions from narrow floating-point products, while FP32 accumulation suffers from quantization losses and costly normalization. To address these limitations, we propose a hybrid precision-scalable reduction tree for MX MACs that combines the benefits of both approaches, enabling efficient mixed-precision accumulation with controlled accuracy relaxation. Moreover, we integrate an 8x8 array of these MACs into the state-of-the-art (SotA) NPU integration platform, SNAX, to provide efficient control and data transfer to our optimized precision-scalable MX datapath. We evaluate our design both on MAC and system level and compare it to the SotA. Our integrated system achieves an energy efficiency of 657, 1438-1675, and 4065 GOPS/W, respectively, for MXINT8, MXFP8/6, and MXFP4, with a throughput of 64, 256, and 512 GOPS.

NPU设计混合精度低功耗计算硬件加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。