arXiv:2603.08977eess.AScs.SD2026-03中稿 · Interspeech 2026被引 1

一种无需大量目标语音即可实现高质量语音转换的通用内容分解方法

Universal Speech Content Factorization

  • 通过最小二乘优化学习通用语音到内容映射,线性可逆提取低秩语音表示
  • 仅需数秒目标语音即可生成高质量语音转换,性能媲美需大量数据的方法
  • 适合需要快速部署、低资源语音转换或音色解耦文本转语音的场景

我们提出通用语音内容分解(USCF),一种简单且可逆的线性方法,用于提取低秩语音表示,其中说话人音色被抑制而语音内容得以保留。USCF 通过最小二乘优化学习通用语音到内容映射,将闭集语音转换方法扩展至开集设置,并仅需少量目标语音即可推导出说话人特异性变换。嵌入分析表明,USCF 能有效去除说话人相关变化。作为零样本语音转换系统,USCF 在可懂度、自然度和说话人相似性方面表现优异,优于需要大量目标语音或额外神经训练的方法。最后,我们证明,作为训练高效的音色解耦语音特征,USCF 特征可作为音色提示式文本转语音模型的声学表示。语音样本与代码已公开。

原文摘要 · Abstract (English)

We propose Universal Speech Content Factorization (USCF), a simple and invertible linear method for extracting a low-rank speech representation in which speaker timbre is suppressed while phonetic content is preserved. USCF extends Speech Content Factorization, a closed-set voice conversion (VC) method, to an open-set setting by learning a universal speech-to-content mapping via least-squares optimization and deriving speaker-specific transformations from only a few seconds of target speech. We show through embedding analysis that USCF effectively removes speaker-dependent variation. As a zero-shot VC system, USCF achieves competitive intelligibility, naturalness, and speaker similarity compared to methods that require substantially more target-speaker data or additional neural training. Finally, we demonstrate that as a training-efficient timbre-disentangled speech feature, USCF features can serve as the acoustic representation for training timbre-prompted text-to-speech models. Speech samples and code are publicly available.

语音转换音色解耦零样本特征提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。