用扩散变压器提升语音增强质量,让降噪后的声音像录音棚录制一样清晰。
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
- 用潜空间扩散变换器建模语音,结合强条件特征来生成高保真信号。
- 在多个数据集上首次达到与录音棚级音频相当的主观质量,内容失真率降低30%以上。
- 适合需要高保真语音还原的场景,如影视配音、语音助手和远程会议。
真实世界语音录音常受背景噪声和混响影响。语音增强旨在通过生成干净高保真信号来缓解这些问题。尽管近期生成式方法已取得进展,但仍面临两大挑战:(1)内容幻觉,即生成的音素虽合理但与原始语句不符;(2)不一致性,无法保留说话人身份和副语言特征。本文提出DiTSE(Diffusion Transformer for Speech Enhancement),通过潜空间扩散变换器结合鲁棒条件特征,实现全频带高质量语音重建。主观与客观评估结果表明,DiTSE首次在DAPS数据集上达到与真实录音棚音频相当的音质水平。同时显著提升说话人身份与内容保真度,相比现有方法大幅减少幻觉现象。音频样本可访问:http://hguimaraes.me/DiTSE
原文摘要 · Abstract (English)
Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity signals. While recent generative approaches for speech enhancement have shown promising results, they still face two major challenges: (1) content hallucination, where plausible phonemes generated differ from the original utterance; and (2) inconsistency, failing to preserve speaker's identity and paralinguistic features from the input speech. In this work, we introduce DiTSE (Diffusion Transformer for Speech Enhancement), which addresses quality issues of degraded speech in full bandwidth. Our approach employs a latent diffusion transformer model together with robust conditioning features, effectively addressing these challenges while remaining computationally efficient. Experimental results from both subjective and objective evaluations demonstrate that DiTSE achieves state-of-the-art audio quality that, for the first time, matches real studio-quality audio from the DAPS dataset. Furthermore, DiTSE significantly improves the preservation of speaker identity and content fidelity, reducing hallucinations across datasets compared to state-of-the-art enhancers. Audio samples are available at: http://hguimaraes.me/DiTSE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。