arXiv:2608.02700cs.LGcs.AI2026-08

针对模拟存内计算噪声,提出自适应精度量化框架,显著提升低比特模型性能。

NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory

论文配图:NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory
图 1 · 摘自论文原文
  • 基于实测噪声建模,动态分配不同精度,避开噪声主导区域的无效高精度
  • 2比特权值量化下,视觉模型准确率提升8.05个百分点,语言模型困惑度降低54.7%
  • 仅用3.2-3.8等效比特即获得接近全精度的收益,适合资源受限的边缘部署

模拟存内计算(CIM)可实现高效神经网络推理,但器件变异和读取噪声会严重损害低比特量化模型性能。现有面向CIM的量化方法主要最小化理想量化误差,忽视硬件噪声底限,导致精度分配低效。本文提出NANQ,一种面向模拟CIM的噪声感知混合精度非均匀量化框架。NANQ基于eFlash CIM阵列实测响应,建模权重幅度相关的噪声特性,并将噪声分布转化为自适应量化密度:在低噪声区域分配更细分辨率,在噪声主导区域避免无效高精度。进一步通过统一阈值识别各层在硬件噪声下的精度饱和点,实现层间位宽分配。片上实验在eFlash CIM SoC上验证,2比特权值幅度量化下,相较于PowerQuant,NANQ使视觉模型准确率提升8.05个百分点,语言模型平均困惑度(PPL)降低54.7%。混合精度的NANQ仅需3.2-3.8等效比特,即可捕获绝大部分额外量化资源带来的增益。

原文摘要 · Abstract (English)

Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade low-bit quantized models. Existing CIM-oriented quantization methods mainly minimize ideal quantization error, ignoring the hardware noise floor and thus causing inefficient precision allocation. We propose NANQ, a noise-aware mixed-precision non-uniform quantization framework for analog CIM. NANQ models magnitude-dependent weight noise from measured responses of an eFlash CIM array and converts the noise profile into an adaptive quantization density, assigning finer resolution to low-noise regions while avoiding ineffective precision in noise-dominated regions. It further assigns layer-wise bit-widths by identifying each layer's precision saturation point under hardware noise using a unified threshold. On-chip experiments on an eFlash CIM SoC show that, under 2-bit weight-magnitude quantization, NANQ improves vision-model accuracy by 8.05 percentage points and reduces language-model PPL by 54.7% on average over PowerQuant. Mixed-precision NANQ captures most of the gains obtainable from additional quantization resources with only 3.2-3.8 equivalent bits.

存内计算量化模拟计算低比特

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。