arXiv:2604.04507cs.ARcs.RO2026-04中稿 · ANRF-sponsored 2nd…

一种支持双精度的低功耗浮点乘加单元,用于高效AI推理。

DHFP-PE: Dual-Precision Hybrid Floating Point Processing Element for AI Acceleration

论文配图:DHFP-PE: Dual-Precision Hybrid Floating Point Processing Element for AI Acceleration
图 1 · 摘自论文原文
  • 用4位单元实现两种精度:单个4×4或两个2×2并行乘法
  • 在28纳米工艺下功耗仅2.13毫瓦,面积缩小60.4%
  • 适合边缘设备的低功耗和混合精度计算场景

人工智能与边缘计算中对低精度算术的快速采用,催生了对节能且灵活的浮点乘加(MAC)单元的强烈需求。本文提出一种支持FP8(E4M3、E5M2)和FP4(2×E2M1、2×E1M2)格式的双精度浮点MAC处理单元,专为低功耗和高吞吐量的AI工作负载优化。该架构采用新型比特分块技术,使单一4位单元乘法器既能作为标准4×4乘法器用于FP8,也能作为两个并行2×2乘法器用于2位操作数,实现硬件利用率最大化,无需重复逻辑。该设计在28纳米工艺下实现1.94 GHz的工作频率,面积为0.00396 mm²,功耗仅为2.13 mW,相比现有最优设计,面积减少60.4%,功耗节省86.6%,非常适合部署于大型加速器架构中的能效受限的AI推理与混合精度计算应用。

原文摘要 · Abstract (English)

The rapid adoption of low-precision arithmetic in artificial intelligence and edge computing has created a strong demand for energy-efficient and flexible floating-point multiply-accumulate (MAC) units. This paper presents a dual-precision floating-point MAC processing element supporting FP8 (E4M3, E5M2) and FP4 (2 x E2M1, 2 x E1M2) formats, specifically optimized for low-power and high-throughput AI workloads. The proposed architecture employs a novel bit-partitioning technique that enables a single 4-bit unit multiplier to operate either as a standard 4 x 4 multiplier for FP8 or as two parallel 2 x 2 multipliers for 2-bit operands, achieving maximum hardware utilization without duplicating logic. Implemented in 28 nm technology, the proposed PE achieves an operating frequency of 1.94 GHz with an area of 0.00396 mm^2 and power consumption of 2.13 mW, resulting in up to 60.4% area reduction and 86.6% power savings compared to state-of-the-art designs, making it well suited for energy-constrained AI inference and mixed-precision computing applications when deployed within larger accelerator architectures.

AI加速低功耗浮点运算硬件设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。