让语音合成同时支持背景消除与保留,且可灵活控制。
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
- 用可控掩码预测策略,让模型决定是否去除背景。
- 在多种声学条件下都能精准控制背景处理,泛化能力强。
- 适合需要灵活处理语音环境的场景,如会议录音、语音助手。
语音中的背景音对自然对话至关重要,提供环境信息帮助理解,但过强的背景会影响语音清晰度。根据情境需求,有时需去除背景以提升可懂度,有时则需保留背景以维持语境完整性。尽管零样本语音合成技术已有进展,现有系统在包含背景音的语音提示下仍表现不佳。为此,我们提出一种可控掩码语音预测策略,结合双说话人编码器,利用任务相关控制信号,引导模型同时实现背景消除与保留。实验表明,该方法可在多种声学条件下精确控制背景处理,并在未见场景中表现出强泛化能力。
原文摘要 · Abstract (English)
The acoustic background plays a crucial role in natural conversation. It provides context and helps listeners understand the environment, but a strong background makes it difficult for listeners to understand spoken words. The appropriate handling of these backgrounds is situation-dependent: Although it may be necessary to remove background to ensure speech clarity, preserving the background is sometimes crucial to maintaining the contextual integrity of the speech. Despite recent advancements in zero-shot Text-to-Speech technologies, current systems often struggle with speech prompts containing backgrounds. To address these challenges, we propose a Controllable Masked Speech Prediction strategy coupled with a dual-speaker encoder, utilizing a task-related control signal to guide the prediction of dual background removal and preservation targets. Experimental results demonstrate that our approach enables precise control over the removal or preservation of background across various acoustic conditions and exhibits strong generalization capabilities in unseen scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。