通过解耦语音因素实现精细可控的语音生成
MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement
- 将语音内容、音色、情感解耦为独立表示,实现精准控制
- 在多因素组合生成任务中,词错误率低至4.67%,主观评分最高
- 适合需要高度可控语音生成的研究者与应用开发者
生成富有表现力且可控制的人类语音是生成式人工智能的核心目标之一,但长期受限于语音因素的深度耦合以及现有控制机制的粗粒度。为此,我们提出了MF-Speech框架,包含两个核心组件:MF-SpeechEncoder和MF-SpeechGenerator。MF-SpeechEncoder作为因子净化器,采用多目标优化策略,将原始语音信号分解为内容、音色和情感等高度纯净且独立的表示。随后,MF-SpeechGenerator作为指挥家,通过动态融合与分层风格自适应归一化(HSAN),实现对这些因子的精确、可组合、细粒度控制。实验表明,在极具挑战性的多因子组合语音生成任务中,MF-Speech显著优于当前最先进方法,词错误率降低至4.67%,风格控制指标SECS=0.5685,相关性Corr=0.68,主观评估得分最高(nMOS=3.96,sMOS_emotion=3.86,sMOS_style=3.78)。此外,学习到的离散因子表现出强迁移能力,显示出作为通用语音表征的巨大潜力。
原文摘要 · Abstract (English)
Generating expressive and controllable human speech is one of the core goals of generative artificial intelligence, but its progress has long been constrained by two fundamental challenges: the deep entanglement of speech factors and the coarse granularity of existing control mechanisms. To overcome these challenges, we have proposed a novel framework called MF-Speech, which consists of two core components: MF-SpeechEncoder and MF-SpeechGenerator. MF-SpeechEncoder acts as a factor purifier, adopting a multi-objective optimization strategy to decompose the original speech signal into highly pure and independent representations of content, timbre, and emotion. Subsequently, MF-SpeechGenerator functions as a conductor, achieving precise, composable and fine-grained control over these factors through dynamic fusion and Hierarchical Style Adaptive Normalization (HSAN). Experiments demonstrate that in the highly challenging multi-factor compositional speech generation task, MF-Speech significantly outperforms current state-of-the-art methods, achieving a lower word error rate (WER=4.67%), superior style control (SECS=0.5685, Corr=0.68), and the highest subjective evaluation scores(nMOS=3.96, sMOS_emotion=3.86, sMOS_style=3.78). Furthermore, the learned discrete factors exhibit strong transferability, demonstrating their significant potential as a general-purpose speech representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。