让语音合成可连续调节声音质感,提升教学与表演应用价值
Speech Synthesis along Perceptual Voice Quality Dimensions
- 用条件连续归一化流实现语音质感的连续调控
- 成功操控粗糙、气息感、共鸣和重量四项声音特质
- 专家评测验证效果,适用于语音训练与创意表达
当前的语音合成或语音转换系统主要关注情绪、口音等抽象韵律特征的控制,而本文聚焦于语音学家识别的感知语音质量(PVQs),这类属性属于更底层的语音特征。能够操控这些质量对语音病理学培训或配音演员具有重要价值。本文将基于条件连续归一化流的方法融入文语转换系统,实现对感知语音属性的连续尺度调节。与以往方法不同,该系统不直接操作声学参数,而是通过示例学习。我们展示了系统在操控粗糙度、气息感、共鸣度和重量四个语音质量维度上的能力。由语音学家对已见及未见说话人条件下的修改结果进行评估,结果揭示了系统的潜力与改进空间。
原文摘要 · Abstract (English)
While expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual voice qualities (PVQs) recognized by phonetic experts, which are speech properties at a lower level of abstraction. The ability to manipulate PVQs can be a valuable tool for teaching speech pathologists in training or voice actors. In this paper, we integrate a Conditional Continuous-Normalizing-Flow-based method into a Text-to-Speech system to modify perceptual voice attributes on a continuous scale. Unlike previous approaches, our system avoids direct manipulation of acoustic correlates and instead learns from examples. We demonstrate the system's capability by manipulating four voice qualities: Roughness, breathiness, resonance and weight. Phonetic experts evaluated these modifications, both for seen and unseen speaker conditions. The results highlight both the system's strengths and areas for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。