arXiv:2409.07556eess.AScs.SD2024-09被引 21

让语音编辑更稳定安全,支持零样本文本控制。

SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

  • 用Transformer解码器+无分类器引导提升生成稳定性。
  • 在RealEdit和LibriTTS任务上达顶尖性能,修复效果优于Encodec。
  • 可检测编辑痕迹,抗背景噪声强,适合真实场景应用。

本文提出SSR-Speech,一种基于Transformer解码器的神经编解码自回归模型,用于实现稳定、安全且鲁棒的零样本文本驱动语音编辑与文生语音合成。通过引入无分类器引导增强生成稳定性,并设计帧级水印编码器(watermark Encodec),在编辑区域嵌入可检测水印,实现编辑痕迹追溯。波形重建利用原始未编辑语音片段,相比Encodec模型提供更优恢复效果。实验表明,SSR-Speech在RealEdit语音编辑任务和LibriTTS文生语音任务中均达到当前最优表现,且在多段落编辑和背景噪声干扰下仍保持显著鲁棒性。代码与演示已公开。

原文摘要 · Abstract (English)

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot textbased speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. The source code and demos are released.

语音编辑零样本生成安全文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。