提出可应对硬件缺陷的量化训练框架,提升低功耗设备上的模型可靠性。
Regularization-based Framework for Quantization-, Fault- and Variability-Aware Training
- 基于正则化设计量化感知训练,支持固定、可学习步长及非均匀量化
- 在CIFAR-10和ImageNet上实现4位量化下竞争力性能,4位SNN在两个数据集表现良好
- 可同时抵御20%位故障和40%器件变异,适合边缘计算与存算一体硬件部署
高效推理对边缘AI设备部署深度学习模型至关重要。3-4位低比特量化结合定点运算可提升效率,而模拟非易失性存储等低功耗内存技术进一步优化。然而,这些方法引入了非理想硬件行为,如位故障和器件间差异。本文提出一种基于正则化的量化感知训练(QAT)框架,支持固定、可学习步长及可学习非均匀量化,在CIFAR-10和ImageNet上取得竞争力结果。该方法扩展至脉冲神经网络(SNNs),在CIFAR10-DVS和N-Caltech 101上实现4位网络优异性能。除量化外,框架还支持故障与变异性感知微调,有效缓解固定位故障(权重位卡死)和器件电阻变异问题。相比先前故障感知训练,本方法在最高20%位故障率和40%器件差异下显著提升性能恢复。结果表明,该框架为量化与鲁棒性协同训练提供通用解决方案,增强低功耗非理想硬件中的效率与可靠性。
原文摘要 · Abstract (English)
Efficient inference is critical for deploying deep learning models on edge AI devices. Low-bit quantization (e.g., 3- and 4-bit) with fixed-point arithmetic improves efficiency, while low-power memory technologies like analog nonvolatile memory enable further gains. However, these methods introduce non-ideal hardware behavior, including bit faults and device-to-device variability. We propose a regularization-based quantization-aware training (QAT) framework that supports fixed, learnable step-size, and learnable non-uniform quantization, achieving competitive results on CIFAR-10 and ImageNet. Our method also extends to Spiking Neural Networks (SNNs), demonstrating strong performance on 4-bit networks on CIFAR10-DVS and N-Caltech 101. Beyond quantization, our framework enables fault and variability-aware fine-tuning, mitigating stuck-at faults (fixed weight bits) and device resistance variability. Compared to prior fault-aware training, our approach significantly improves performance recovery under upto 20% bit-fault rate and 40% device-to-device variability. Our results establish a generalizable framework for quantization and robustness-aware training, enhancing efficiency and reliability in low-power, non-ideal hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。