用轻量模型实现语音宽带扩展,提升低码率传输质量
A lightweight and robust method for blind wideband-to-fullband extension of speech
- 基于传统语音编码思路,设计仅370K参数的轻量扩展模型
- 在6~12kb/s下显著提升语音质量,9kb/s时达3GPP EVS水平
- 无需额外信号即可实现高质量扩展,适合现有编码器升级
在低带宽或低复杂度场景中,压缩语音带宽是常见做法。本文提出一种轻量且鲁棒的宽带语音信号扩展方法,灵感源自经典语音编码技术。模型仅需约370K参数和约140 MFLOPS(或70 MMACS)计算量,帧长10ms,前视仅0.27ms,适用于主流宽带语音编码器。通过与Opus SILK编码器(1.5版)结合,在P.808 DCR听感测试中验证了其在6~12 kb/s下显著提升质量。同时证明,使用该扩展方法的Opus 1.5在9 kb/s时达到3GPP EVS 9.6 kb/s及Opus 1.4 18 kb/s的音质水平,表明盲式带宽扩展可媲美传统引导式扩展,为向后兼容的质量提升提供路径。
原文摘要 · Abstract (English)
Reducing the bandwidth of speech is common practice in resource constrained environments like low-bandwidth speech transmission or low-complexity vocoding. We propose a lightweight and robust method for extending the bandwidth of wideband speech signals that is inspired by classical methods developed in the speech coding context. The resulting model has just ~370K parameters and a complexity of ~140 MFLOPS (or ~70 MMACS). With a frame size of 10 ms and a lookahead of only 0.27 ms, the model is well-suited for use with common wideband speech codecs. We evaluate the model's robustness by pairing it with the Opus SILK speech codec (1.5 release) and verify in a P.808 DCR listening test that it significantly improves quality from 6 to 12 kb/s. We also demonstrate that Opus 1.5 together with the proposed bandwidth extension at 9 kb/s meets the quality of 3GPP EVS at 9.6 kb/s and that of Opus 1.4 at 18 kb/s showing that the blind bandwidth extension can meet the quality of classical guided bandwidth extensions thus providing a way for backward-compatible quality improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。