arXiv:2503.20999cs.SDcs.GR2025-03

用文本控制语音风格,实现更自然的声线转换

Text-Driven Voice Conversion via Latent State-Space Modeling

  • 将语音视为连续隐空间中的动态系统,分离内容与风格
  • 通过文本提示精准控制语速、重音和说话人特征
  • 相比现有方法,风格切换更平滑,伪影更少

文本驱动的语音转换可基于文本描述定制说话人特征与韵律。然而,现有方法多依赖直接的文本到语音训练,难以灵活控制细微风格或音色特征。本文提出一种新型潜空间状态模型语音转换方法(LSS-VC),将每段语音视为连续隐空间中的演化动力系统。受Mamba启发——该模型曾用于高效文本驱动图像风格迁移——我们将其适配至语音风格转换。具体而言,学习一个语音隐流形,使风格与内容可通过文本风格提示独立操控。设计自适应跨模态融合机制,将风格信息注入语音隐表示,实现对说话人身份、语速和强调的可解释、细粒度控制。大量实验表明,该方法在主观与客观评价指标上均显著优于近期基线,且风格间过渡更平滑、伪影更少、文本控制更精确。

原文摘要 · Abstract (English)

Text-driven voice conversion allows customization of speaker characteristics and prosodic elements using textual descriptions. However, most existing methods rely heavily on direct text-to-speech training, limiting their flexibility in controlling nuanced style elements or timbral features. In this paper, we propose a novel \textbf{Latent State-Space} approach for text-driven voice conversion (\textbf{LSS-VC}). Our method treats each utterance as an evolving dynamical system in a continuous latent space. Drawing inspiration from mamba, which introduced a state-space model for efficient text-driven \emph{image} style transfer, we adapt a loosely related methodology for \emph{voice} style transformation. Specifically, we learn a voice latent manifold where style and content can be manipulated independently by textual style prompts. We propose an adaptive cross-modal fusion mechanism to inject style information into the voice latent representation, enabling interpretable and fine-grained control over speaker identity, speaking rate, and emphasis. Extensive experiments show that our approach significantly outperforms recent baselines in both subjective and objective quality metrics, while offering smoother transitions between styles, reduced artifacts, and more precise text-based style control.

语音转换隐空间文本控制风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。