融合文本与音素级自监督特征,提升语音转换的保真度与语义完整性。
FabasedVC: Enhancing Voice Conversion with Text Modality Fusion and Phoneme-Level SSL Features
- 用文本编码器融合音素、声调和BERT特征,增强内容表征
- 通过平均池化与注意力机制将帧级自监督特征转为音素级特征
- 引入时长预测器,精准对齐目标说话人语速与韵律,适合语音克隆场景
在语音转换(VC)中,保持完整语义信息的同时准确建模目标说话人的音色与语调至关重要。本文提出FabasedVC,一种基于VITS的端到端语音转换系统,通过融合文本模态信息、音素级自监督学习(SSL)特征以及时长预测器,实现与目标说话人更一致的音色、语调和时长,并提升内容完整性。具体地,采用文本特征编码器对文本、音素、声调及BERT特征进行编码;利用两种方法将帧级SSL特征转化为音素级特征:平均池化与基于音素持续时间的注意力机制;同时引入时长预测器以更好对齐目标说话人的语速与韵律。实验结果表明,该方法在自然度、相似性与内容完整性方面均优于现有系统。
原文摘要 · Abstract (English)
In voice conversion (VC), it is crucial to preserve complete semantic information while accurately modeling the target speaker's timbre and prosody. This paper proposes FabasedVC to achieve VC with enhanced similarity in timbre, prosody, and duration to the target speaker, as well as improved content integrity. It is an end-to-end VITS-based VC system that integrates relevant textual modality information, phoneme-level self-supervised learning (SSL) features, and a duration predictor. Specifically, we employ a text feature encoder to encode attributes such as text, phonemes, tones and BERT features. We then process the frame-level SSL features into phoneme-level features using two methods: average pooling and attention mechanism based on each phoneme's duration. Moreover, a duration predictor is incorporated to better align the speech rate and prosody of the target speaker. Experimental results demonstrate that our method outperforms competing systems in terms of naturalness, similarity, and content integrity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。