arXiv:2508.21524cs.ARcs.LG2025-08被引 5

提出二值权重多比特激活量化方法,提升存内计算加速器的能效与精度。

Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators

  • 采用闭式解优化每层二值权重,增强二值化表示能力。
  • 在CIFAR-10和ImageNet上分别提升1.44%–5.46%和0.35%–5.37%准确率。
  • 4比特激活量化实现硬件开销与模型性能的最佳平衡,适合边缘部署。

存内计算(CIM)加速器已成为提升卷积神经网络(CNN)能效的有前途方案。将CNN部署于CIM平台通常需对网络权重和激活进行量化以满足硬件约束。然而,现有方法或侧重硬件效率,采用二值权重与激活量化而牺牲精度;或使用多比特权重与激活以提升精度,但效率受限。本文提出一种面向基于CIM加速器的二值权重多比特激活(BWMA)方法。主要贡献包括:推导各层权重量化的闭式解,显著提升二值权重的表征能力;设计可微函数实现激活量化,近似理想多比特函数,避免繁琐的最优参数搜索。在CIFAR-10和ImageNet数据集上的全面实验表明,该方法相比现有方法显著提升精度,分别取得1.44%–5.46%和0.35%–5.37%的提升。此外,硬件仿真结果显示,4比特激活量化在硬件成本与模型性能间达到最佳平衡。

原文摘要 · Abstract (English)

Compute-in-memory (CIM) accelerators have emerged as a promising way for enhancing the energy efficiency of convolutional neural networks (CNNs). Deploying CNNs on CIM platforms generally requires quantization of network weights and activations to meet hardware constraints. However, existing approaches either prioritize hardware efficiency with binary weight and activation quantization at the cost of accuracy, or utilize multi-bit weights and activations for greater accuracy but limited efficiency. In this paper, we introduce a novel binary weight multi-bit activation (BWMA) method for CNNs on CIM-based accelerators. Our contributions include: deriving closed-form solutions for weight quantization in each layer, significantly improving the representational capabilities of binarized weights; and developing a differentiable function for activation quantization, approximating the ideal multi-bit function while bypassing the extensive search for optimal settings. Through comprehensive experiments on CIFAR-10 and ImageNet datasets, we show that BWMA achieves notable accuracy improvements over existing methods, registering gains of 1.44\%-5.46\% and 0.35\%-5.37\% on respective datasets. Moreover, hardware simulation results indicate that 4-bit activation quantization strikes the optimal balance between hardware cost and model performance.

存内计算量化轻量级模型边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。