无需配对数据,用多任务学习实现更清晰的语音转换。
Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning
- 设计双流输入结构,强化说话人与内容的解耦
- 自毁式修正约束提升内容编码器性能
- 支持非平行数据,训练成本低且音质高
语音转换(VC)在保留语言内容的同时改变声音特征。本文提出Stepback网络,一种基于非平行数据的新型语音转换模型。不同于依赖平行数据的传统方法,该方法利用深度学习技术增强解耦效果并更好保留语言内容。Stepback网络采用不同域数据的双流输入,并引入自毁式修正约束优化内容编码器。大量实验表明,该模型显著提升语音转换性能,降低训练成本,同时实现高质量语音转换。其设计为高级语音转换任务提供了有前景的解决方案。
原文摘要 · Abstract (English)
Voice conversion (VC) modifies voice characteristics while preserving linguistic content. This paper presents the Stepback network, a novel model for converting speaker identity using non-parallel data. Unlike traditional VC methods that rely on parallel data, our approach leverages deep learning techniques to enhance disentanglement completion and linguistic content preservation. The Stepback network incorporates a dual flow of different domain data inputs and uses constraints with self-destructive amendments to optimize the content encoder. Extensive experiments show that our model significantly improves VC performance, reducing training costs while achieving high-quality voice conversion. The Stepback network's design offers a promising solution for advanced voice conversion tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。