提出可抗噪的零样本语音编辑框架,提升真实环境下的语音生成质量。
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
- 基于频段感知的降噪模块与上下文增强策略
- 在含噪场景下优于现有方法,保持语音清晰度与自然度
- 适合实际应用中存在背景噪声的语音编辑任务
随着零样本文本到语音技术的快速发展,生成的语音信号已几乎无法与真实语音区分。语音编辑(如插入和替换)因其潜在应用价值受到研究者关注。然而,现有研究仅考虑纯净语音场景。在真实应用中,环境噪声会显著降低生成质量。本文提出一种抗噪语音编辑框架 SeamlessEdit,采用频段感知降噪模块与上下文增强策略,有效应对语音与噪声频段重叠的情况。所提框架在多个定量与定性评估中均优于当前最优方法。
原文摘要 · Abstract (English)
With the fast development of zero-shot text-to-speech technologies, it is possible to generate high-quality speech signals that are indistinguishable from the real ones. Speech editing, including speech insertion and replacement, appeals to researchers due to its potential applications. However, existing studies only considered clean speech scenarios. In real-world applications, the existence of environmental noise could significantly degrade the quality of generation. In this study, we propose a noise-resilient speech editing framework, SeamlessEdit, for noisy speech editing. SeamlessEdit adopts a frequency-band-aware noise suppression module and an in-content refinement strategy. It can well address the scenario where the frequency bands of voice and background noise are not separated. The proposed SeamlessEdit framework outperforms state-of-the-art approaches in multiple quantitative and qualitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。