通过语义对齐减少音色泄漏,实现更自然的零样本语音转换
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
- 利用文本与音频表征对齐,解耦语音内容与说话人特征
- 在无显式说话人嵌入下实现高保真转换,音色相似度提升显著
- 适合需要隐私保护和跨说话人泛化的语音合成应用
零样本语音转换(VC)在保持语言和副语言内容的同时,将语音转换为目标说话人的音色。然而,在神经编解码器和大语言模型(LLM)基的VC中,量化表示会混淆说话人身份与内容,导致音色泄漏问题。本文提出SemAlignVC,采用一种名为SemAlign的新方法,对齐文本与音频表征,确保说话人无关的语义编码。该解耦表征作为自回归变压器的条件,实现高保真转换,且无需显式说话人嵌入。实验表明,SemAlignVC显著降低音色泄漏,在音色相似度、可懂性和自然度上优于基线模型,是一种鲁棒、隐私友好且具备强泛化能力的语音转换方案。音频样例可访问:https://shivammehta25.github.io/SemAlignVC/
原文摘要 · Abstract (English)
Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker traits persist-remains a challenge, especially in neural codec and LLM-based VC, where quantized representations entangle speaker identity with content. We introduce SemAlignVC, an architecture designed to prevent timbre leakage using SemAlign, a novel method that aligns text and audio representations to ensure speaker-independent semantic encoding. This disentangled representation conditions an autoregressive transformer for high-fidelity conversion without explicit speaker embeddings. Experiments show SemAlignVC significantly reduces timbre leakage, outperforming baselines in speaker timbre similarity, intelligibility, and naturalness, making it a robust, privacy-preserving, and generalizable VC solution. Audio samples can be accessed at https://shivammehta25.github.io/SemAlignVC/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。