arXiv:2511.05143eess.AS2025-11被引 1

通过流形变换实现对嘶哑音质的精准合成控制。

Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice

  • 用归一化流构建全局声学属性调节模块,实现对嘶哑音质的非局部控制。
  • 主观测试表明合成语音质量略降(MOS评分下降),但嘶哑音质操控成功。
  • 适合语音教学、语音病理研究等需精确控制声学特征的场景。

在文本转语音(TTS)系统中,对感知语音特质进行可控合成对教学应用具有重要意义,尤其在难以直观理解的语音现象展示方面。本文提出一种基于归一化流的全局说话人属性调控模块,可有效操纵非持续性、局部化的嘶哑音质,避免依赖通常不可靠的逐帧嘶哑检测器。主观听觉测试显示,经调控的语音在音质上成功实现了嘶哑效果,其平均意见分(MOS)相比原始录音略有下降,但仍保持可接受水平。该方法为语音合成中的精细声学控制提供了新路径。

原文摘要 · Abstract (English)

The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustrate pho- netic concepts that are otherwise difficult to grasp. Here, we show that a TTS system, that is augmented with a global speaker attribute manipulation block based on normalizing flows1 , is capable of correctly manipulating the non-persistent, localized quality of creaky voice, thus avoiding the necessity of a, typi- cally unreliable, frame-wise creak predictor. Subjective listen- ing tests confirm successful creak manipulation at a slightly re- duced MOS score compared to the original recording.

语音合成声学控制嘶哑音质归一化流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。