arXiv:2507.06079cs.LGcs.AI2025-07被引 1

量化训练让结构化状态空间模型更高效,适配边缘硬件部署。

QS4D: Quantization-aware training for efficient hardware deployment of structured state-space sequential models

  • 通过量化感知训练降低模型复杂度,最多减少两个数量级。
  • 验证了模型规模与精度的关系,提升对模拟噪声的鲁棒性。
  • 适用于内存计算芯片,显著提升边缘设备算力效率。

结构化状态空间模型(SSM)是一类新兴的深度学习模型,特别适合处理长序列数据。其恒定的内存开销相较于线性增长内存需求的Transformer模型,使其成为资源受限边缘计算设备的理想选择。尽管已有研究探讨量化感知训练(QAT)对SSM的影响,但通常未考虑其在专用边缘硬件(如模拟存内计算,AIMC芯片)上的应用。本文表明,QAT可使SSM在多个性能指标上复杂度降低高达两个数量级。我们分析了模型大小与数值精度之间的关系,发现QAT增强了对模拟噪声的鲁棒性,并支持结构化剪枝。最后,我们将这些技术整合,成功将SSM部署于忆阻型模拟存内计算硬件,显著提升了计算效率。

原文摘要 · Abstract (English)

Structured State Space models (SSM) have recently emerged as a new class of deep learning models, particularly well-suited for processing long sequences. Their constant memory footprint, in contrast to the linearly scaling memory demands of Transformers, makes them attractive candidates for deployment on resource-constrained edge-computing devices. While recent works have explored the effect of quantization-aware training (QAT) on SSMs, they typically do not address its implications for specialized edge hardware, for example, analog in-memory computing (AIMC) chips. In this work, we demonstrate that QAT can significantly reduce the complexity of SSMs by up to two orders of magnitude across various performance metrics. We analyze the relation between model size and numerical precision, and show that QAT enhances robustness to analog noise and enables structural pruning. Finally, we integrate these techniques to deploy SSMs on a memristive analog in-memory computing substrate and highlight the resulting benefits in terms of computational efficiency.

状态空间模型量化训练边缘计算存内计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。