arXiv:2511.23203cs.ARcs.AI2025-11

通过灵活降压提升低精度神经网络加速能效,误差几乎不变。

GAVINA: flexible aggressive undervolting for bit-serial mixed-precision DNN acceleration

  • 选择性降低最低有效位的电压,结合串行计算实现灵活近似。
  • 在最激进配置下达到89 TOP/sW能效,比基准提升20%。
  • 适合追求极致能效的低精度深度学习硬件设计者。

电压超调(即降压)是实现节能型深度神经网络(DNN)加速的诱人近似技术,因功耗与电压呈平方关系。然而其极高的错误率阻碍了广泛应用。此外,现有降压加速器依赖8比特计算,无法与先进的低精度(<8比特)架构竞争。为此,我们提出一种新方法——受保护的激进降压(GAV),融合降压与位串行计算思想,在少数最低有效位组合上激进降低供电电压。基于此,我们实现了GAVINA(GAV 混合精度加速器),支持任意混合精度和灵活降压,最激进配置下能效高达89 TOP/sW。通过建立GAVINA的误差模型,我们证明:通过降压可实现20%的能效提升,且在ResNet-18上精度损失可忽略。

原文摘要 · Abstract (English)

Voltage overscaling, or undervolting, is an enticing approximate technique in the context of energy-efficient Deep Neural Network (DNN) acceleration, given the quadratic relationship between power and voltage. Nevertheless, its very high error rate has thwarted its general adoption. Moreover, recent undervolting accelerators rely on 8-bit arithmetic and cannot compete with state-of-the-art low-precision (<8b) architectures. To overcome these issues, we propose a new technique called Guarded Aggressive underVolting (GAV), which combines the ideas of undervolting and bit-serial computation to create a flexible approximation method based on aggressively lowering the supply voltage on a select number of least significant bit combinations. Based on this idea, we implement GAVINA (GAV mIxed-precisioN Accelerator), a novel architecture that supports arbitrary mixed precision and flexible undervolting, with an energy efficiency of up to 89 TOP/sW in its most aggressive configuration. By developing an error model of GAVINA, we show that GAV can achieve an energy efficiency boost of 20% via undervolting, with negligible accuracy degradation on ResNet-18.

低精度计算能效优化硬件加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。