分离音高与声带振动,让语音合成保留说话人身份
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
- 用条件连续归一化流增强训练数据,实现音高与声带振动解耦
- 在多种声带振动强度下,说话人识别准确率显著提升
- 适合需保持身份一致的语音克隆与个性化合成场景
我们提出一种系统,能够在忠实改变声带振动感知质量的同时,保留说话人的身份感知。尽管高声带振动概率通常与低音高相关,但这种关联仅在群体层面成立,并非普遍适用。通过在语音合成系统的训练数据中引入基于条件连续归一化流的说话人操控模块,实现了音高与声带振动的解耦。实验表明,在不同声带振动强度下,该系统显著提升了说话人验证性能。
原文摘要 · Abstract (English)
We introduce a system capable of faithfully modifying the perceptual voice quality of creak while preserving the speaker's perceived identity. While it is well known that high creak probability is typically correlated with low pitch, it is important to note that this is a property observed on a population of speakers but does not necessarily hold across all situations. Disentanglement of pitch from creak is achieved by augmentation of the training dataset of a speech synthesis system with a speaker manipulation block based on conditional continuous normalizing flow. The experiments show greatly improved speaker verification performance over a range of creak manipulation strengths.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。