用语义一致性降低扩散模型语音增强的生成伪影,提升准确率。
ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement
- 通过语音嵌入方差预测发音错误,指导推理过程
- 多轮扩散推理使低信噪比下字错率降低15%
- 自适应扩散步数平衡去伪影与推理延迟
基于扩散模型的语音增强(SE)虽能生成自然语音并具备强泛化能力,但存在生成伪影和高推理延迟等关键问题。本文系统研究了扩散模型语音增强中的伪影预测与抑制方法。研究表明,语音嵌入的方差可用于推断推理过程中的发音错误。基于此,我们提出一种由多个扩散运行间语义一致性引导的集成推理方法,在低信噪比条件下将字错率(WER)降低15%,显著提升发音准确性和语义合理性。最后,我们分析了扩散步数的影响,发现自适应扩散步数可在抑制伪影与控制延迟之间取得平衡。研究结果表明,语义先验是引导生成式语音增强实现无伪影输出的强大工具。
原文摘要 · Abstract (English)
Diffusion-based speech enhancement (SE) achieves natural-sounding speech and strong generalization, yet suffers from key limitations like generative artifacts and high inference latency. In this work, we systematically study artifact prediction and reduction in diffusion-based SE. We show that variance in speech embeddings can be used to predict phonetic errors during inference. Building on these findings, we propose an ensemble inference method guided by semantic consistency across multiple diffusion runs. This technique reduces WER by 15% in low-SNR conditions, effectively improving phonetic accuracy and semantic plausibility. Finally, we analyze the effect of the number of diffusion steps, showing that adaptive diffusion steps balance artifact suppression and latency. Our findings highlight semantic priors as a powerful tool to guide generative SE toward artifact-free outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。