用流模型在编码器潜空间高效恢复语音高频信息,提升清晰度。
CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent Space
- 在神经音频编码器的紧凑潜空间中使用条件流模型重建语音
- 8kHz→16kHz和44.1kHz任务上实现高保真频谱与优质听感
- 适合语音增强、通信系统等对带宽效率敏感的应用
语音带宽扩展通过恢复或推断低带宽语音中的高频内容,提升清晰度与可懂性。现有方法多依赖频谱或波形建模,计算开销大且高频保真度有限。神经音频编码器提供紧凑的潜空间表示,更完整保留声学细节,但因表征不匹配,精确恢复高分辨率潜变量仍具挑战。本文提出CodecFlow,一种基于神经编码器的带宽扩展框架,在紧凑潜空间中实现高效语音重建。该方法在连续编码器嵌入上采用语音活动感知的条件流转换器,并引入结构约束残差向量量化器以增强潜空间对齐稳定性。全程端到端优化,可在8 kHz → 16 kHz及44.1 kHz语音带宽扩展任务中实现优异的频谱保真度与感知质量。
原文摘要 · Abstract (English)
Speech Bandwidth Extension improves clarity and intelligibility by restoring/inferring appropriate high-frequency content for low-bandwidth speech. Existing methods often rely on spectrogram or waveform modeling, which can incur higher computational cost and have limited high-frequency fidelity. Neural audio codecs offer compact latent representations that better preserve acoustic detail, yet accurately recovering high-resolution latent information remains challenging due to representation mismatch. We present CodecFlow, a neural codec-based BWE framework that performs efficient speech reconstruction in a compact latent space. CodecFlow employs a voicing-aware conditional flow converter on continuous codec embeddings and a structure-constrained residual vector quantizer to improve latent alignment stability. Optimized end-to-end, CodecFlow achieves strong spectral fidelity and enhanced perceptual quality on 8 kHz to 16 kHz and 44.1 kHz speech BWE tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。