用语义通信生成语音,比传统方法更保真。
Generative Semantic Communication for Text-to-Speech Synthesis
- 用预训练模型和向量量化构建收发端语义知识库
- 融合Transformer与扩散模型实现高效语义编码
- 在噪声和衰落信道下均显著优于四种基线
语义通信通过仅传输源数据的语义信息,有望提升通信效率。然而,传统语义通信方法主要面向数据重建任务,对文本转语音(TTS)等生成式任务效率不高。为此,本文提出一种面向TTS合成的新型生成式语义通信框架,利用生成式人工智能技术。首先,采用预训练大模型WavLM与残差向量量化方法,在发送端和接收端分别构建两个语义知识库(KB),发送端KB用于有效提取语义,接收端KB用于实现逼真语音合成。其次,采用Transformer编码器与扩散模型实现高效语义编码,通信开销低。最后的数值结果表明,该框架在加性高斯白噪声信道和瑞利衰落信道下,生成语音的保真度显著高于四种基线方法。
原文摘要 · Abstract (English)
Semantic communication is a promising technology to improve communication efficiency by transmitting only the semantic information of the source data. However, traditional semantic communication methods primarily focus on data reconstruction tasks, which may not be efficient for emerging generative tasks such as text-to-speech (TTS) synthesis. To address this limitation, this paper develops a novel generative semantic communication framework for TTS synthesis, leveraging generative artificial intelligence technologies. Firstly, we utilize a pre-trained large speech model called WavLM and the residual vector quantization method to construct two semantic knowledge bases (KBs) at the transmitter and receiver, respectively. The KB at the transmitter enables effective semantic extraction, while the KB at the receiver facilitates lifelike speech synthesis. Then, we employ a transformer encoder and a diffusion model to achieve efficient semantic coding without introducing significant communication overhead. Finally, numerical results demonstrate that our framework achieves much higher fidelity for the generated speech than four baselines, in both cases with additive white Gaussian noise channel and Rayleigh fading channel.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。