一通搞定多种语音失真修复,48kHz高清语音秒复原
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
- 用单一潜空间生成流程,统一处理各类语音修复任务
- 通过联合神经先验降低模型负担,实现无蒸馏一步修复
- 在真实与合成语音上均表现卓越,适配多场景语音增强
桥接模型已被用于语音增强,但多为单任务,通用语音修复(GSR)能力受限。本文提出VoiceBridge,一种一步式潜空间桥接模型(LBM),可高效从多种失真中重建48 kHz全频带语音。为继承数据域桥接模型优势,设计了能量保持的变分自编码器,提升波形-潜空间在不同能量水平下的对齐性。通过将波形压缩为连续潜表示,VoiceBridge以单一潜空间到潜空间生成过程建模多种GSR任务,基于可扩展的Transformer架构。为缓解从差异显著的低质量先验重构高质量目标的挑战,提出联合神经先验,统一减轻模型在多样化任务中的负担。进一步通过联合优化LBM、解码器与判别器,调整桥接训练目标,使模型由降噪器转变为生成器,实现无需蒸馏的一步式通用语音修复。在域内(如去噪、超分辨率)与域外任务(如合成语音优化)及多个数据集上的广泛验证表明,VoiceBridge性能优越。
原文摘要 · Abstract (English)
Bridge models have been investigated in speech enhancement but are mostly single-task, with constrained general speech restoration (GSR) capability. In this work, we propose VoiceBridge, a one-step latent bridge model (LBM) for GSR, capable of efficiently reconstructing 48 kHz fullband speech from diverse distortions. To inherit the advantages of data-domain bridge models, we design an energy-preserving variational autoencoder, enhancing the waveform-latent space alignment over varying energy levels. By compressing waveform into continuous latent representations, VoiceBridge models~\textit{various} GSR tasks with a~\textit{single} latent-to-latent generative process backed by a scalable transformer. To alleviate the challenge of reconstructing the high-quality target from distinctively different low-quality priors, we propose a joint neural prior for GSR, uniformly reducing the burden of the LBM in diverse tasks. Building upon these designs, we further investigate bridge training objective by jointly tuning LBM, decoder and discriminator together, transforming the model from a denoiser to generator and enabling \textit{one-step GSR without distillation}. Extensive validation across in-domain (\textit{e.g.}, denoising and super-resolution) and out-of-domain tasks (\textit{e.g.}, refining synthesized speech) and datasets demonstrates the superior performance of VoiceBridge. Demos: https://VoiceBridgedemo.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。