通过可证明的误差上界压缩深度状态空间模型,60%参数量减少不损失性能
A Deep State-Space Model Compression Method using Upper Bound on Output Error
- 基于输出误差上界设计层间误差约束,以h²范数指导压缩
- 在IMDb任务上实现60%参数量缩减,无需重新训练
- 适合需要轻量化部署且要求性能稳定的深度状态空间模型用户
我们研究包含线性二次输出(LQO)系统作为内部模块的深度状态空间模型(Deep SSMs),提出一种具有可证明输出误差保证的压缩方法。首先推导出两个Deep SSMs之间输出误差的上界,并表明该上界可表示为各层LQO系统之间的h²-误差范数之和。特别地,我们发现减小浅层LQO系统的h²逼近误差能有效降低输出误差上界。随后,针对该上界构建优化问题,并开发基于梯度的模型降阶方法。在LRA基准的IMDb任务上的数值实验验证了所提方法的有效性:在不重新训练的情况下,可将可训练参数减少约60%,同时保持原始模型性能。
原文摘要 · Abstract (English)
We study deep state-space models (Deep SSMs) that contain linear quadratic-output (LQO) systems as internal blocks and present a compression method with a provable output error guarantee. We first derive an upper bound on the output error between two Deep SSMs and show that the bound can be expressed in terms of the $h^2$-error norms between the layerwise LQO systems. In particular, we show that reducing the $h^2$ approximation errors of the LQO systems placed in shallow layers is effective in reducing the derived upper bound on the output error. Next, we formulate an optimization problem for the derived upper bound and develop a gradient-based MOR method. In the numerical experiments, using the IMDb task from the LRA benchmark, we demonstrate the effectiveness of the proposed upper-bound-based compression method. In particular, we show that the number of trainable parameters can be reduced by approximately 60\% without retraining while maintaining the performance of the original model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。