QUADS统一优化量化与蒸馏,实现低资源下的高效语音理解。
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
- 联合优化量化与蒸馏,分阶段训练提升压缩效率。
- 在SLURP上达71.13%准确率,计算量降低60–73倍。
- 适合部署于边缘设备的语音识别系统,性能稳定。
语音语言理解(SLU)系统需在性能与效率间取得平衡,尤其在资源受限环境下。现有方法分别应用蒸馏与量化,导致压缩效果不佳,因蒸馏未考虑量化约束。本文提出QUADS框架,通过预调优模型进行多阶段训练,统一优化两者,在低比特条件下保持高适应性与准确性。QUADS在SLURP数据集上达到71.13%准确率,在FSC数据集上达99.20%,相比先进模型仅下降最多5.56%。同时,计算复杂度降低60–73倍(以GMACs计),模型大小缩小83–700倍,展现出极强的极端量化鲁棒性。该成果为真实场景中资源受限的SLU应用提供了高效解决方案。
原文摘要 · Abstract (English)
Spoken Language Understanding (SLU) systems must balance performance and efficiency, particularly in resource-constrained environments. Existing methods apply distillation and quantization separately, leading to suboptimal compression as distillation ignores quantization constraints. We propose QUADS, a unified framework that optimizes both through multi-stage training with a pre-tuned model, enhancing adaptability to low-bit regimes while maintaining accuracy. QUADS achieves 71.13\% accuracy on SLURP and 99.20\% on FSC, with only minor degradations of up to 5.56\% compared to state-of-the-art models. Additionally, it reduces computational complexity by 60--73$\times$ (GMACs) and model size by 83--700$\times$, demonstrating strong robustness under extreme quantization. These results establish QUADS as a highly efficient solution for real-world, resource-constrained SLU applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。