arXiv:2609.03381eess.AS2026-09

轻量级流式语音超分辨率模型,实时处理无延迟。

StreamWSR: Streamable and Lightweight Waveform-Domain Neural Speech Super-Resolution

论文配图:StreamWSR: Streamable and Lightweight Waveform-Domain Neural Speech Super-Resolution
图 1 · 摘自论文原文
  • 全因果结构+帧级波形表示,支持零前瞻流式推理。
  • 仅900万参数、20亿浮点运算,性能优于主流方法。
  • 适合低延迟语音增强场景,如实时通话与语音助手。

本文提出 StreamWSR,一种可流式处理的波形域语音超分辨率模型。采用全因果架构与紧凑的帧级波形表示,实现零前瞻流式推理,无需声码器重建或显式相位预测。StreamWSR 通过步长因果卷积将输入波形下采样为紧凑的帧级表示,再利用轻量级因果长短时建模主干,在因果约束下捕捉局部波形结构与长程历史依赖。最后,通过因果转置卷积将建模输出还原至波形域,并通过残差连接与输入波形融合,生成最终高分辨率语音。在16 kHz语音超分辨率任务上,实验表明,StreamWSR 在语音质量与可懂度方面达到或超越代表性波形与谱域基线方法,同时保持零前瞻流式优势,仅需900万参数和20亿浮点运算。

原文摘要 · Abstract (English)

This paper proposes StreamWSR, a Streamable neural Waveform-domain model for speech Super-Resolution (SR). By adopting a fully causal architecture with compact frame-level waveform representation, the proposed StreamWSR supports zero-look-ahead streaming inference while avoiding vocoder-based reconstruction and explicit phase prediction. Specifically, StreamWSR downsamples the input waveform into a compact frame-level representation using strided causal convolutions. Then, a lightweight causal long-short-term modeling backbone is employed to capture both local waveform structures and long-range historical dependencies under causal constraints. Finally, the modeled output is converted back to the waveform domain through a causal transposed-convolution and combined with the input waveform via a residual connection to generate the final high-resolution speech. Experimental results on 16 kHz speech SR show that StreamWSR achieves competitive or superior speech quality and intelligibility compared with representative waveform- and spectrum-based baselines, while maintaining a zero-look-ahead streaming advantage with only 9M parameters and 2G FLOPs.

语音增强流式处理波形模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。