arXiv:2412.08312cs.SDcs.LG2024-12

统一模型实现语音与歌声的音色与口音转换,保持语调自然。

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction

  • 基于HuBERT编码器与HiFi-GAN解码器,融合音高与歌手嵌入
  • 可对语音与歌声混合样本进行口音转换,保留原内容与韵律
  • 适用于配音、内容创作及TTS/IVR系统,提升语音风格灵活性

本文提出一种新型语音转换模型,可实现语音与歌声的双向转换。针对现有系统在情感传递、发音与口音变化、非语言声音重现等方面的挑战,该模型具备处理语音与歌唱混合样本的能力,可在不改变原内容和语调的前提下实现口音转换。模型采用基于HuBERT的编码器提取声学与语言特征,配合HiFi-GAN解码器生成目标说话人音色。通过引入基频(f0)特征与歌手嵌入,有效保障转换过程中音高、音调与声线身份的准确性。该方法显著提升了语音风格转换的自然度与灵活性,展现出在语音配音、内容创作以及文本转语音(TTS)和交互式语音应答(IVR)系统中的广泛应用潜力。

原文摘要 · Abstract (English)

This paper presents a new voice conversion model capable of transforming both speaking and singing voices. It addresses key challenges in current systems, such as conveying emotions, managing pronunciation and accent changes, and reproducing non-verbal sounds. One of the model's standout features is its ability to perform accent conversion on hybrid voice samples that encompass both speech and singing, allowing it to change the speaker's accent while preserving the original content and prosody. The proposed model uses an encoder-decoder architecture: the encoder is based on HuBERT to process the speech's acoustic and linguistic content, while the HiFi-GAN decoder audio matches the target speaker's voice. The model incorporates fundamental frequency (f0) features and singer embeddings to enhance performance while ensuring the pitch & tone accuracy and vocal identity are preserved during transformation. This approach improves how naturally and flexibly voice style can be transformed, showing strong potential for applications in voice dubbing, content creation, and technologies like Text-to-Speech (TTS) and Interactive Voice Response (IVR) systems.

语音转换口音转换歌声合成自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。