无需配对数据即可控制语音中的年龄与性别,同时保留说话人身份。
Controlling your Attributes in Voice
- 用GAN-VAE分离说话人身份与属性特征
- 两阶段语音转换实现自然的年龄/性别变化
- 适合语音合成与个性化语音应用
语音生成中的属性控制旨在修改年龄、性别等个人属性,同时保持源语音的身份信息。尽管图像生成中面部属性控制已取得进展,但语音生成领域的类似方法仍不充分。本文提出一种无需平行数据的说话人属性控制新方法:首先通过基于GAN的说话人表示变分自编码器从说话人向量中提取身份与属性信息;再利用两阶段语音转换模型捕捉语音中自然的属性表达。实验表明,该方法不仅在说话人表示层面实现属性控制,还能在语音层面改变年龄与性别,同时保持语音质量和说话人身份不变。
原文摘要 · Abstract (English)
Attribute control in generative tasks aims to modify personal attributes, such as age and gender while preserving the identity information in the source sample. Although significant progress has been made in controlling facial attributes in image generation, similar approaches for speech generation remain largely unexplored. This letter proposes a novel method for controlling speaker attributes in speech without parallel data. Our approach consists of two main components: a GAN-based speaker representation variational autoencoder that extracts speaker identity and attributes from speaker vector, and a two-stage voice conversion model that captures the natural expression of speaker attributes in speech. Experimental results show that our proposed method not only achieves attribute control at the speaker representation level but also enables manipulation of the speaker age and gender at the speech level while preserving speech quality and speaker identity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。