用倒放语音增强说话人表征,提升语音转换保真度。
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
- 利用倒放语音学习说话人特征作为数据增强。
- 显著提升说话人相似度指标,保持高语音质量。
- 适合需要精准保留音色的语音转换场景。
语音时间倒放是指将整个语音信号在时间上反转,使其反向播放。此类信号因音素和音节结构被破坏而完全无法理解,但仍保留语调模式,使听者可感知说话人身份。本文提出利用从倒放语音中学习到的说话人表征作为数据增强策略,以提升语音转换中的说话人表征能力。在最先进的基于扩散模型的语音转换框架中评估该方法,实验结果表明,该方法显著提升了与说话人相似度相关的评分,同时保持了高质量的语音输出。
原文摘要 · Abstract (English)
Speech time reversal refers to the process of reversing the entire speech signal in time, causing it to play backward. Such signals are completely unintelligible since the fundamental structures of phonemes and syllables are destroyed. However, they still retain tonal patterns that enable perceptual speaker identification despite losing linguistic content. In this paper, we propose leveraging speaker representations learned from time reversed speech as an augmentation strategy to enhance speaker representation. Notably, speaker and language disentanglement in voice conversion (VC) is essential to accurately preserve a speaker's unique vocal traits while minimizing interference from linguistic content. The effectiveness of the proposed approach is evaluated in the context of state-of-the-art diffusion-based VC models. Experimental results indicate that the proposed approach significantly improves speaker similarity-related scores while maintaining high speech quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。