通过平均特征解耦语音内容与音色,提升语音转换质量。
AVENet: Disentangling Features by Approximating Average Features for Voice Conversion
- 用对齐语音数据的平均特征作为理想内容表示。
- AVENet能有效生成接近平均特征的解耦表征。
- 适合追求音色与内容分离的语音转换研究者。
语音转换(VC)在特征解耦方面取得进展,但仍难以平衡音色与内容信息。本文评估了语音转换中常用预训练模型的特征表现,并提出一种创新的语音特征解耦方法。具体而言,首先定义了一种理想内容特征——平均特征,通过帧级对齐平行语音(FAPS)数据计算得到。为生成FAPS数据,采用冻结文语转换系统中的时长预测器并调整说话人嵌入的技术。为使该平均特征适配传统VC数据集,设计了AVENet,以输入特征生成逼近平均特征的输出。在VC系统中测试了AVENet提取特征的性能,实验结果表明其优于多种现有语音特征解耦方法,验证了该解耦策略的有效性。
原文摘要 · Abstract (English)
Voice conversion (VC) has made progress in feature disentanglement, but it is still difficult to balance timbre and content information. This paper evaluates the pre-trained model features commonly used in voice conversion, and proposes an innovative method for disentangling speech feature representations. Specifically, we first propose an ideal content feature, referred to as the average feature, which is calculated by averaging the features within frame-level aligned parallel speech (FAPS) data. For generating FAPS data, we utilize a technique that involves freezing the duration predictor in a Text-to-Speech system and manipulating speaker embedding. To fit the average feature on traditional VC datasets, we then design the AVENet to take features as input and generate closely matching average features. Experiments are conducted on the performance of AVENet-extracted features within a VC system. The experimental results demonstrate its superiority over multiple current speech feature disentangling methods. These findings affirm the effectiveness of our disentanglement approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。