arXiv:2409.09098cs.SDcs.CL2024-09中稿 · ICASSP 2025被引 22

让语音合成能零样本生成高保真口音,支持从未见过的口音。

AccentBox: Towards High-Fidelity Zero-Shot Accent Generation

  • 两阶段流程:先识别口音,再用口音嵌入控制语音合成。
  • 未见说话人上口音识别达0.56 F1,生成口音更自然。
  • 可生成新口音,适合多口音语音应用开发。

尽管近期零样本文本转语音(ZS-TTS)模型在自然度和说话人相似性方面表现优异,但在口音保真度与控制力上仍有不足。为此,本文提出统一外语口音转换(FAC)、带口音语音合成与零样本语音合成的零样本口音生成方法,采用新颖的两阶段流程。第一阶段中,基于预训练的口音识别模型,在未见说话人上达到0.56的F1分数,优于现有方法。第二阶段,将零样本语音合成系统以该模型提取的无说话人依赖口音嵌入作为条件,显著提升固有口音与跨口音生成的保真度,并实现对未曾见过口音的生成能力。

原文摘要 · Abstract (English)

While recent Zero-Shot Text-to-Speech (ZS-TTS) models have achieved high naturalness and speaker similarity, they fall short in accent fidelity and control. To address this issue, we propose zero-shot accent generation that unifies Foreign Accent Conversion (FAC), accented TTS, and ZS-TTS, with a novel two-stage pipeline. In the first stage, we achieve state-of-the-art (SOTA) on Accent Identification (AID) with 0.56 f1 score on unseen speakers. In the second stage, we condition a ZS-TTS system on the pretrained speaker-agnostic accent embeddings extracted by the AID model. The proposed system achieves higher accent fidelity on inherent/cross accent generation, and enables unseen accent generation.

语音合成口音生成零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。