语音与背景音生成系统在双赛道中分别获第二、第一,实现零样本语音风格克隆。
The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge
- 用Single-Codec将语音离散化,结合语言模型实现零样本语音风格迁移。
- Track 1和Track 2分别获得第二名和第一名,48kHz高保真音频合成成功。
- 基于大模型生成场景适配的背景音描述,融合语音与伴奏效果自然。
本文介绍提交至ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge(ICAGC)的NPU-HWC系统。该系统包含两个模块:用于Track 1的语音生成器和用于Track 2的背景音频生成器。在Track 1中,采用Single-Codec将语音转为离散标记,通过基于语言模型的方法实现零样本说话风格克隆。Single-Codec在标记层有效解耦音色与语调特征,减轻自回归语言模型的声学建模负担。同时,使用DSPGAN将16 kHz mel-spectrograms上采样至48 kHz高保真波形。在Track 2中,提出一种基于大语言模型(LLM)的背景音频生成系统,可生成符合场景的伴奏描述,利用Tango 2合成背景音频,并与Track 1生成的语音融合。最终提交结果在两赛道中分别获得第二名和第一名。
原文摘要 · Abstract (English)
This paper presents the NPU-HWC system submitted to the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC). Our system consists of two modules: a speech generator for Track 1 and a background audio generator for Track 2. In Track 1, we employ Single-Codec to tokenize the speech into discrete tokens and use a language-model-based approach to achieve zero-shot speaking style cloning. The Single-Codec effectively decouples timbre and speaking style at the token level, reducing the acoustic modeling burden on the autoregressive language model. Additionally, we use DSPGAN to upsample 16 kHz mel-spectrograms to high-fidelity 48 kHz waveforms. In Track 2, we propose a background audio generator based on large language models (LLMs). This system produces scene-appropriate accompaniment descriptions, synthesizes background audio with Tango 2, and integrates it with the speech generated by our Track 1 system. Our submission achieves the second place and the first place in Track 1 and Track 2 respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。