自动匹配语音风格,让合成语音更自然生动。
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
- 用检索增强生成技术动态选择最适语音风格。
- 在多个数据集上实现95%以上风格匹配准确率。
- 适合需要个性化语音合成的场景如虚拟助手。
随着语音合成技术的进步,用户对合成语音的自然度和表现力提出了更高要求。然而以往研究忽视了提示选择的重要性。本文提出一种基于检索增强生成(RAG)技术的文本到语音(TTS)框架,可依据文本内容动态调整语音风格,实现更自然、生动的表达效果。我们构建了一个包含多种语境下高质量语音样本的语音风格知识库,并设计了一套风格匹配方案。该方案利用Llama、PER-LLM-Embedder和Moka提取的嵌入向量,与知识库中的样本进行匹配,选出最适合当前文本的语音风格用于合成。此外,实证研究验证了该方法的有效性。演示视频可访问:https://thuhcsi.github.io/icme2025-AutoStyle-TTS
原文摘要 · Abstract (English)
With the advancement of speech synthesis technology, users have higher expectations for the naturalness and expressiveness of synthesized speech. But previous research ignores the importance of prompt selection. This study proposes a text-to-speech (TTS) framework based on Retrieval-Augmented Generation (RAG) technology, which can dynamically adjust the speech style according to the text content to achieve more natural and vivid communication effects. We have constructed a speech style knowledge database containing high-quality speech samples in various contexts and developed a style matching scheme. This scheme uses embeddings, extracted by Llama, PER-LLM-Embedder,and Moka, to match with samples in the knowledge database, selecting the most appropriate speech style for synthesis. Furthermore, our empirical research validates the effectiveness of the proposed method. Our demo can be viewed at: https://thuhcsi.github.io/icme2025-AutoStyle-TTS
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。