通过掩码学习卷积核的半结构化稀疏模式,实现推理加速且不损失性能。
Inducing Semi-Structured Sparsity by Masking for Efficient Model Inference in Convolutional Networks
- 用掩码学习卷积核的半结构化稀疏模式,适配现有硬件加速。
- 推理速度提升两倍以上,模型性能无下降。
- 掩码对预测影响可量化,支持模型更新后的稳定性保证。
卷积模型在视觉任务及基础模型骨干中扮演关键角色,亟需高效加速技术。本文提出一种新方法,通过掩码学习卷积核的半结构化稀疏模式,以利用现有硬件加速。该方法使卷积模型推理速度提升两倍以上,同时保持原始模型权重和结构不变,便于后续更新。此外,掩码对预测的影响可量化,推导出在模型更新后仍能保持预测稳定性的理论保证。
原文摘要 · Abstract (English)
The crucial role of convolutional models, both as standalone vision models and backbones in foundation models, necessitates effective acceleration techniques. This paper proposes a novel method to learn semi-structured sparsity patterns for convolution kernels in the form of maskings enabling the utilization of readily available hardware accelerations. The approach accelerates convolutional models more than two-fold during inference without decreasing model performance. At the same time, the original model weights and structure remain unchanged keeping the model thus easily updatable. Beyond the immediate practical use, the effect of maskings on prediction is easily quantifiable. Therefore, guarantees on model predictions under maskings are derived showing stability bounds for learned maskings even after updating the original underlying model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。