arXiv:2506.02414cs.MMcs.CL2025-06中稿 · Interspeech 2025, …被引 2

StarVC统一建模文本与语音生成,提升语音转换保真度。

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

  • 先预测文本再生成声学特征,实现端到端联合建模。
  • 在语音转换中同时降低词错误率(WER)和音素错误率(CER)。
  • 适合需要高保真语音转换的研究者或语音合成系统开发者。

语音转换(VC)在保留语言内容的前提下改变语音以匹配目标说话人。传统方法通常直接从语音中提取说话人信息,忽视了语言内容的显式利用。由于VC本质上需解耦说话人身份与语言内容,引入结构化语义特征可提升转换效果。然而,先前将语义特征融入VC的尝试效果有限,促使我们引入显式文本建模。本文提出StarVC,一种统一的自回归语音转换框架,先预测文本标记,再合成声学特征。实验表明,StarVC在保持语言内容(即词错误率WER和音素错误率CER)和说话人特征(即说话人相似度评分SECS和主观评价得分MOS)方面均优于传统方法。音频演示可访问:https://thuhcsi.github.io/StarVC/

原文摘要 · Abstract (English)

Voice Conversion (VC) modifies speech to match a target speaker while preserving linguistic content. Traditional methods usually extract speaker information directly from speech while neglecting the explicit utilization of linguistic content. Since VC fundamentally involves disentangling speaker identity from linguistic content, leveraging structured semantic features could enhance conversion performance. However, previous attempts to incorporate semantic features into VC have shown limited effectiveness, motivating the integration of explicit text modeling. We propose StarVC, a unified autoregressive VC framework that first predicts text tokens before synthesizing acoustic features. The experiments demonstrate that StarVC outperforms conventional VC methods in preserving both linguistic content (i.e., WER and CER) and speaker characteristics (i.e., SECS and MOS). Audio demo can be found at: https://thuhcsi.github.io/StarVC/.

语音转换自回归模型文本生成联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。