arXiv:2509.25842cs.AI2025-09被引 3

通过分层预测风格嵌入,让语音合成更精准地响应文本指令。

HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis

  • 分两阶段预测风格嵌入,先按音色再按风格细分。
  • 相比传统方法,风格控制更精准,语音自然度不变。
  • 结合统计与人耳偏好生成更准确的提示词,适合语音合成研究者。

可控语音合成可通过调节语调、音量、语速、音高及音高波动等属性实现对说话风格的精确控制。随着大语言模型和扩散模型的发展,可控文本到语音系统已从标签驱动转向自然语言描述驱动,通常通过从文本提示中预测全局风格嵌入来实现。然而,这种简单预测忽略了风格嵌入的内在分布,限制了系统的潜力。本研究通过 t-SNE 分析主流 TTS 系统的风格嵌入分布,发现其呈现清晰的层级聚类:先按音色分组,再按风格属性细分为子簇。基于此,提出 HiStyle——一种两阶段条件风格嵌入预测器,并引入对比学习对齐文本与音频嵌入空间。此外,提出一种融合统计方法与人耳偏好的风格标注策略,生成更准确且感知一致的文本提示。大量实验表明,应用于基础 TTS 模型时,HiStyle 在风格可控性上显著优于其他方法,同时保持高语音自然度与可懂度。音频样本见 https://anonymous.4open.science/w/HiStyle-2517/。

原文摘要 · Abstract (English)

Controllable speech synthesis refers to the precise control of speaking style by manipulating specific prosodic and paralinguistic attributes, such as gender, volume, speech rate, pitch, and pitch fluctuation. With the integration of advanced generative models, particularly large language models (LLMs) and diffusion models, controllable text-to-speech (TTS) systems have increasingly transitioned from label-based control to natural language description-based control, which is typically implemented by predicting global style embeddings from textual prompts. However, this straightforward prediction overlooks the underlying distribution of the style embeddings, which may hinder the full potential of controllable TTS systems. In this study, we use t-SNE analysis to visualize and analyze the global style embedding distribution of various mainstream TTS systems, revealing a clear hierarchical clustering pattern: embeddings first cluster by timbre and subsequently subdivide into finer clusters based on style attributes. Based on this observation, we propose HiStyle, a two-stage style embedding predictor that hierarchically predicts style embeddings conditioned on textual prompts, and further incorporate contrastive learning to help align the text and audio embedding spaces. Additionally, we propose a style annotation strategy that leverages the complementary strengths of statistical methodologies and human auditory preferences to generate more accurate and perceptually consistent textual prompts for style control. Comprehensive experiments demonstrate that when applied to the base TTS model, HiStyle achieves significantly better style controllability than alternative style embedding predicting approaches while preserving high speech quality in terms of naturalness and intelligibility. Audio samples are available at https://anonymous.4open.science/w/HiStyle-2517/.

语音合成风格控制嵌入预测文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。