arXiv:2412.16977eess.AS2024-12中稿 · ICASSP 2025被引 3

让合成语音保持环境特征,支持未见说话人零样本语音合成。

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis

  • 分步解耦环境、说话人与文本信息,逐步提取关键特征。
  • 在多个指标上优于现有方法,环境与说话人相似度均达领先水平。
  • 适合需要环境感知的语音合成与转换任务,如虚拟助手场景。

本文提出一种基于增量解耦的环境感知零样本语音合成方法IDEA-TTS,可在不依赖目标说话人数据的情况下,合成具有指定环境特征的语音。该方法以VITS为骨干网络,设计增量解耦流程:先通过环境估计算法将环境声谱分解为环境掩码和增强声谱;环境掩码经环境编码器生成环境嵌入,增强声谱则结合预训练的环境鲁棒说话人编码器提取的说话人嵌入,实现说话人与文本因子的解耦。最终,说话人与环境嵌入共同条件化解码器,完成环境感知语音生成。实验表明,IDEA-TTS在语音质量、说话人相似度与环境相似度方面均表现优异,且在声学环境转换任务中达到当前最佳性能。

原文摘要 · Abstract (English)

This paper proposes an Incremental Disentanglement-based Environment-Aware zero-shot text-to-speech (TTS) method, dubbed IDEA-TTS, that can synthesize speech for unseen speakers while preserving the acoustic characteristics of a given environment reference speech. IDEA-TTS adopts VITS as the TTS backbone. To effectively disentangle the environment, speaker, and text factors, we propose an incremental disentanglement process, where an environment estimator is designed to first decompose the environmental spectrogram into an environment mask and an enhanced spectrogram. The environment mask is then processed by an environment encoder to extract environment embeddings, while the enhanced spectrogram facilitates the subsequent disentanglement of the speaker and text factors with the condition of the speaker embeddings, which are extracted from the environmental speech using a pretrained environment-robust speaker encoder. Finally, both the speaker and environment embeddings are conditioned into the decoder for environment-aware speech generation. Experimental results demonstrate that IDEA-TTS achieves superior performance in the environment-aware TTS task, excelling in speech quality, speaker similarity, and environmental similarity. Additionally, IDEA-TTS is also capable of the acoustic environment conversion task and achieves state-of-the-art performance.

语音合成零样本环境感知解耦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。