UniPASE统一处理多采样率语音增强,低幻觉高保真。
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations

- 基于知识蒸馏的统一表征模块,直接转换受损波形为清晰语音表征。
- 在多个数据集上性能优于或媲美现有最优模型,支持16kHz到48kHz多采样率。
- 适合需要高保真语音修复且避免语义错乱的工业级语音增强场景。
通用语音增强(USE)旨在从多种失真和不同采样率中恢复语音信号。我们提出UniPASE,是针对USE优化的低幻觉PASE框架扩展。核心是DeWavLM-Omni,一个从WavLM通过大规模监督多失真数据集知识蒸馏微调的统一表征级增强模块。该模块直接将退化波形转换为清晰且语言忠实的语音表征,确保鲁棒增强并最小化语义幻觉。基于这些增强的语音表征,适配器生成包含丰富声学细节的声学表征,再由神经声码器重建对应高保真16 kHz波形。后置网络将波形转换至48~kHz,再重采样回原始采样率,实现多采样率输入输出的无缝处理。在多个评估数据集上的实验结果表明,UniPASE在子任务与全任务上均达到优于或媲美现有最先进模型的性能。该模型也是我们参加URGENT 2026挑战赛的主干模型,在客观评价中获得第一名。源代码与音频演示见https://github.com/xiaobin-rong/unipase/。
原文摘要 · Abstract (English)
Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPASE, an extension of the low-hallucination PASE framework tailored for USE. At its core is DeWavLM-Omni, a unified representation-level enhancement module fine-tuned from WavLM via knowledge distillation on a large-scale supervised multi-distortion dataset. This module directly converts degraded waveforms into clean and linguistically faithful phonetic representations, ensuring robust enhancement with minimal linguistic hallucination. Based on these enhanced phonetic representations, an Adapter generates enhanced acoustic representations containing rich acoustic details, which a neural Vocoder uses to reconstruct corresponding high-fidelity 16-kHz waveforms. A PostNet then converts the waveforms to 48~kHz before resampling them to their original rates, enabling seamless handling of inputs and outputs at multiple sampling rates. Experimental results on several evaluation datasets, covering sub-tasks and full tasks, demonstrate that UniPASE achieves superior or competitive performance compared with existing state-of-the-art models. The proposed model also serves as the backbone of our submission to the URGENT 2026 Challenge, which achieved 1st place in the objective evaluation. The source code and audio demos are available at https://github.com/xiaobin-rong/unipase/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。