StuPASE提升语音增强质量,低幻觉下实现录音室级音质。
StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
- 用无混响干信号微调PASE,显著改善去混响效果
- 替换GAN生成模块为流匹配,强噪声下仍达录音室级音质
- 适合追求高保真且防幻觉的语音增强应用
在生成式语音增强(SE)中,实现高感知质量且无幻觉仍是挑战。代表性方法PASE虽抗幻觉能力强,但在恶劣条件下感知质量有限。本文提出StuPASE,基于PASE实现录音室级质量的同时保持低幻觉特性。首先,我们发现使用干信号目标而非含模拟早期反射的目标进行微调,能显著提升去混响性能。其次,为解决强加性噪声下的性能瓶颈,将PASE中的GAN生成模块替换为流匹配模块,使模型在极端挑战条件下仍可生成高质量语音。实验表明,StuPASE持续产出高感知质量语音,同时保持低幻觉,优于当前最先进方法。音频演示见:https://xiaobin-rong.github.io/stupase_demo/。
原文摘要 · Abstract (English)
Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE). A representative approach, PASE, is robust to hallucination but has limited perceptual quality under adverse conditions. We propose StuPASE, built upon PASE to achieve studio-level quality while retaining its low-hallucination property. First, we show that finetuning PASE with dry targets rather than targets containing simulated early reflections substantially improves dereverberation. Second, to address performance limitations under strong additive noise, we replace the GAN-based generative module in PASE with a flow-matching module, enabling studio-quality generation even under highly challenging conditions. Experiments demonstrate that StuPASE consistently produces perceptually high-quality speech while maintaining low hallucination, outperforming state-of-the-art SE methods. Audio demos are available at: https://xiaobin-rong.github.io/stupase_demo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。