arXiv:2603.08737cs.ARcs.AI2026-03被引 3

通过敏感度引导压缩,实现存算加速器的高效低精度部署。

Sensitivity-Guided Framework for Pruned and Quantized Reservoir Computing Accelerators

  • 基于敏感度识别可删减权重,动态平衡精度与效率。
  • 4位量化+15%剪枝使资源占用降1.2%,功耗延迟积降50.8%。
  • 适合需要高能效的时序数据硬件加速场景。

本文提出一种针对储层计算的压缩框架,支持系统性探索量化精度、剪枝率、模型精度与硬件效率之间的权衡关系。该方法利用敏感度驱动的剪枝机制,识别并移除对模型精度影响较小的量化权重,从而降低计算开销并保持精度。我们通过三个时序数据集(包含分类与回归任务)进行广泛评估,验证了该框架的有效性及剪枝与量化对性能和硬件参数的影响。实验结果表明,在FPGA实现中,该方法在保持高精度的同时显著提升计算与资源效率。例如,在MELBOEN数据集上,4位量化配合15%剪枝可使资源利用率降低1.2%,功率延迟积减少50.8%,且精度无明显下降。

原文摘要 · Abstract (English)

This paper presents a compression framework for Reservoir Computing that enables systematic design-space exploration of trade-offs among quantization levels, pruning rates, model accuracy, and hardware efficiency. The proposed approach leverages a sensitivity-based pruning mechanism to identify and remove less critical quantized weights with minimal impact on model accuracy, thereby reducing computational overhead while preserving accuracy. We perform an extensive trade-off analysis to validate the effectiveness of the proposed framework and the impact of pruning and quantization on model performance and hardware parameters. For this evaluation, we employ three time-series datasets, including both classification and regression tasks. Experimental results across selected benchmarks demonstrate that our proposed approach maintains high accuracy while substantially improving computational and resource efficiency in FPGA-based implementations, with variations observed across different configurations and time series applications. For instance, for the MELBOEN dataset, an accelerator quantized to 4-bit at a 15\% pruning rate reduces resource utilization by 1.2\% and the Power Delay Product (PDP) by 50.8\% compared to an unpruned model, without any noticeable degradation in accuracy.

储层计算量化剪枝FPGA加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。