arXiv:2508.14004cs.LGcs.IT2025-08被引 3

提出渐进式量化方法,让低比特神经网络在极简位宽下仍保持高精度。

GDNSQ: Gradual Differentiable Noise Scale Quantization for Low-bit Neural Networks

  • 用可学习的噪声尺度和位宽,实现全程可微的量化训练。
  • 在W1A1极端设置下仍保持竞争力准确率,优于传统方法。
  • 适合需要极致压缩的边缘设备部署场景。

量化神经网络可视为一系列带噪信道,每层舍入操作随位宽降低而损失容量,浮点检查点决定输入速率上限。我们追踪平均位宽下降时的容量动态,将微调建模为平滑约束优化问题,识别量化瓶颈。方法采用全可微的直通估计器(STE),引入可学习的位宽、噪声尺度与截断边界,并通过外点惩罚强制目标位宽;适度的度量平滑(通过蒸馏)稳定训练过程。尽管结构简单,该方法在极端的W1A1设置下仍达到竞争性精度,同时保持STE的高效性。

原文摘要 · Abstract (English)

Quantized neural networks can be viewed as a chain of noisy channels, where rounding in each layer reduces capacity as bit-width shrinks; the floating-point (FP) checkpoint sets the maximum input rate. We track capacity dynamics as the average bit-width decreases and identify resulting quantization bottlenecks by casting fine-tuning as a smooth, constrained optimization problem. Our approach employs a fully differentiable Straight-Through Estimator (STE) with learnable bit-width, noise scale and clamp bounds, and enforces a target bit-width via an exterior-point penalty; mild metric smoothing (via distillation) stabilizes training. Despite its simplicity, the method attains competitive accuracy down to the extreme W1A1 setting while retaining the efficiency of STE.

低比特量化可微量化神经网络压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。