arXiv:2412.03074cs.CLcs.SD2024-12被引 1

用自监督模型生成的离散符号合成语音,更保真音色语调。

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model

  • 用SSL模型提取语音离散符号代替文本进行语音合成。
  • 离散符号比文本更擅长保留语调、语气等声学特征。
  • 适合研究无需文本标注的语音合成技术者阅读。

本文通过分析使用自监督学习(SSL)模型从原始音频中获得的无文本语音表示,探究了以这些表示替代传统文本表示进行语音合成的效果。由于原始音频缺乏与之配对的语音表示(如转写文本),从非配对语音中获取语音表示对于扩充语音合成数据集至关重要。具体而言,该研究采用SSL模型生成的离散符号表示进行语音合成,并对其合成效果进行了分析。实验结果表明:使用文本表示有助于保持语义信息,而使用离散符号表示在保留声学内容(包括韵律和语调)方面表现更优。

原文摘要 · Abstract (English)

We examine the text-free speech representations of raw audio obtained from a self-supervised learning (SSL) model by analyzing the synthesized speech using the SSL representations instead of conventional text representations. Since raw audio does not have paired speech representations as transcribed texts do, obtaining speech representations from unpaired speech is crucial for augmenting available datasets for speech synthesis. Specifically, the proposed speech synthesis is conducted using discrete symbol representations from the SSL model in comparison with text representations, and analytical examinations of the synthesized speech have been carried out. The results empirically show that using text representations is advantageous for preserving semantic information, while using discrete symbol representations is superior for preserving acoustic content, including prosodic and intonational information.

语音合成自监督学习无文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。