通过遮蔽语音单元提升语音转换中的身份与语义解耦效果
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
- 在编码前遮蔽与音素高度相关的离散语音单元,降低说话人特征对语音内容的依赖
- 在多类语音转换方法中提升解耦效果,注意力模型下客观可懂度提升44%
- 输入级设计通用性强,适配各类编码器-解码器架构的语音转换系统
语音转换(VC)旨在保留语言内容的同时改变说话人身份。现有方法多采用编码器-解码器结构,其解耦性能受限于说话人特征对语音内容的依赖,尤其在基于注意力的方法中更为显著。为此,本文提出一种新型输入级遮蔽机制,在说话人编码前对与音素类别高度相关的离散语音单元进行遮蔽,以限制语音内容对说话人特征的影响。该方法不依赖特定模型结构,可广泛应用于各类编码器-解码器型语音转换框架。实验表明,该方法显著提升了多类语音转换模型的解耦能力,尤其在注意力机制中表现突出,客观可懂度相对提升44%。
原文摘要 · Abstract (English)
Voice conversion (VC) aims to modify the speaker's identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling the speaker's identity from linguistic information is crucial. However, the disentanglement approaches used in these methods are limited as the speaker features depend on the phonetic content of the utterance, compromising disentanglement. This dependency is amplified with attention-based methods. To address this, we introduce a novel masking mechanism in the input before speaker encoding, masking certain discrete speech units that correspond highly with phoneme classes. Our work aims to reduce the phonetic dependency of speaker features by restricting access to some phonetic information. Furthermore, since our approach is at the input level, it is applicable to any encoder-decoder based VC framework. Our approach improves disentanglement and conversion performance across multiple VC methods, showing significant effectiveness, particularly in attention-based method, with 44% relative improvement in objective intelligibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。