通过帧重采样与子带剪枝,降低语音增强计算开销66%以上。
Lightweight Front-end Enhancement for Robust ASR via Frame Resampling and Sub-Band Pruning
- 逐层帧重采样结合残差连接,减少信息损失。
- 逐步剪除低信息量频带,计算量降低超66%。
- 适合部署在资源受限的实时语音识别场景。
近年来自动语音识别(ASR)取得显著进展,但在噪声环境下的鲁棒性仍具挑战。尽管语音增强(SE)前端广泛用于降噪预处理,但通常带来不可忽视的计算开销。本文提出优化方案,在不损害ASR性能的前提下降低SE计算成本。方法结合逐层帧重采样与渐进式子带剪枝:帧重采样在层内下采样输入,利用残差连接缓解信息丢失;子带剪枝逐步剔除低信息量频带,进一步降低计算需求。在合成及真实噪声数据集上的大量实验表明,该系统相比标准BSRNN,SE计算开销降低超过66%,同时保持优异的ASR表现。
原文摘要 · Abstract (English)
Recent advancements in automatic speech recognition (ASR) have achieved notable progress, whereas robustness in noisy environments remains challenging. While speech enhancement (SE) front-ends are widely used to mitigate noise as a preprocessing step for ASR, they often introduce computational non-negligible overhead. This paper proposes optimizations to reduce SE computational costs without compromising ASR performance. Our approach integrates layer-wise frame resampling and progressive sub-band pruning. Frame resampling downsamples inputs within layers, utilizing residual connections to mitigate information loss. Simultaneously, sub-band pruning progressively excludes less informative frequency bands, further reducing computational demands. Extensive experiments on synthetic and real-world noisy datasets demonstrate that our system reduces SE computational overhead over 66 compared to the standard BSRNN, while maintaining strong ASR performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。