让语音转换同时控制音色和环境声学特征,无需成对数据即可精准调节。
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
- 通过文本输入分别指定目标音色与环境,实现独立控制。
- 在合成数据上训练,保留源内容且音色/环境匹配度高。
- 适合需要灵活生成不同场景语音的应用,如虚拟角色配音。
我们提出 TES-VC(Text-driven Environment and Speaker controllable Voice Conversion),一种可独立控制说话人音色与环境声学特征的文本驱动语音转换框架。该方法同时接收目标音色与环境的文本描述,能准确生成符合描述的语音,同时保持源语音内容不变。基于潜在扩散建模,在合成数据上训练并解耦声学与音色特征,有效消除属性间干扰。引入基于检索的音色控制(RBTC)模块,仅需抽象文本描述即可实现精准操控,无需成对数据。实验表明,TES-VC 在音色与环境匹配度、内容保留率及可控性方面均表现优异,具备广泛应用场景潜力。
原文摘要 · Abstract (English)
We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and environmental acoustics. TES-VC processes simultaneous text inputs for target voice and environment, accurately generating speech matching described timbre/environment while preserving source content. Trained on synthetic data with decoupled vocal/environment features via latent diffusion modeling, our method eliminates interference between attributes. The Retrieval-Based Timbre Control (RBTC) module enables precise manipulation using abstract descriptions without paired data. Experiments confirm TES-VC effectively generates contextually appropriate speech in both timbre and environment with high content retention and superior controllability which demonstrates its potential for widespread applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。