arXiv:2505.21245cs.SDeess.AS2025-05中稿 · Interspeech2025被引 5

实现1比特语音识别模型压缩,性能几乎无损失。

Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision

  • 用多精度协同训练和随机精度提升低比特量化效果。
  • 在Switchboard和LibriSpeech数据集上实现1比特无损压缩。
  • 适合边缘设备部署的超低资源语音识别系统。

随着现代语音系统规模迅速增长,模型压缩成为迫切需求。本文研究模型权重量化,直接降低内存占用以适配计算资源受限的应用场景。提出新方法,通过多精度模型协同训练、随机精度机制及张量级可学习缩放因子,缓解低比特量化带来的性能损失。所提方法可在300小时Switchboard和960小时LibriSpeech语料上训练的Conformer自动语音识别系统中,实现2比特与1比特无性能损失量化。最大整体压缩比分别达16.2倍和16.6倍,且词错误率(WER)与全精度基线系统相比无统计显著差异。

原文摘要 · Abstract (English)

Model compression has become an emerging need as the sizes of modern speech systems rapidly increase. In this paper, we study model weight quantization, which directly reduces the memory footprint to accommodate computationally resource-constrained applications. We propose novel approaches to perform extremely low-bit (i.e., 2-bit and 1-bit) quantization of Conformer automatic speech recognition systems using multiple precision model co-training, stochastic precision, and tensor-wise learnable scaling factors to alleviate quantization incurred performance loss. The proposed methods can achieve performance-lossless 2-bit and 1-bit quantization of Conformer ASR systems trained with the 300-hr Switchboard and 960-hr LibriSpeech corpus. Maximum overall performance-lossless compression ratios of 16.2 and 16.6 times are achieved without a statistically significant increase in the word error rate (WER) over the full precision baseline systems, respectively.

语音识别模型压缩低比特量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。