用扩散模型修复长达1秒的语音缺失,保留说话人特征和环境音
Transient Noise Removal via Diffusion-based Speech Inpainting
- 基于扩散模型的语音补全框架,采用音素级分类器引导增强重建质量
- 可处理最长1秒的语音空白,即使在强瞬态噪声下仍保持身份与语调一致
- 无需文本输入也能有效工作,适合真实场景如烟花、敲门等噪声干扰
本文提出PGDI,一种基于扩散模型的语音补全框架,用于恢复缺失或严重损坏的语音片段。不同于以往方法在说话人差异或长间隔下的表现不佳,PGDI能准确重建长达1秒的语音空白,同时保留说话人身份、语调及混响等环境因素。核心在于音素级分类器引导,显著提升重建保真度。该方法为说话人无关设计,在长段完全被强瞬态噪声遮蔽时仍具鲁棒性,适用于真实场景如烟花、门撞击、锤击和建筑噪音。通过在多样化说话人和不同间隙长度下的大量实验,验证了PGDI在复杂声学条件下的优越补全性能。我们还评估了推理时是否有字幕信息:虽有文本可进一步提升效果,但无文本时模型依然有效。
原文摘要 · Abstract (English)
In this paper, we present PGDI, a diffusion-based speech inpainting framework for restoring missing or severely corrupted speech segments. Unlike previous methods that struggle with speaker variability or long gap lengths, PGDI can accurately reconstruct gaps of up to one second in length while preserving speaker identity, prosody, and environmental factors such as reverberation. Central to this approach is classifier guidance, specifically phoneme-level guidance, which substantially improves reconstruction fidelity. PGDI operates in a speaker-independent manner and maintains robustness even when long segments are completely masked by strong transient noise, making it well-suited for real-world applications, such as fireworks, door slams, hammer strikes, and construction noise. Through extensive experiments across diverse speakers and gap lengths, we demonstrate PGDI's superior inpainting performance and its ability to handle challenging acoustic conditions. We consider both scenarios, with and without access to the transcript during inference, showing that while the availability of text further enhances performance, the model remains effective even in its absence. For audio samples, visit: https://mordehaym.github.io/PGDI/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。