arXiv:2412.06602cs.CLcs.AI2024-12EMNLP综述被引 52

系统梳理大模型时代语音合成的可控性技术进展

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

  • 按架构、控制策略和特征表示分类梳理方法
  • 涵盖从传统控制到自然语言提示的全链条技术
  • 适合语音合成与大模型交叉研究者参考

文本到语音(TTS)已从生成自然语音发展到对情感、音色、风格等属性实现细粒度控制。受工业需求增长和深度学习突破(如扩散模型与大语言模型)推动,可控语音合成成为快速发展的研究领域。本综述首次全面回顾可控语音合成方法,涵盖传统控制技术与基于自然语言提示的新兴方法。我们对模型架构、控制策略、特征表示进行分类,并总结该领域的挑战、常用数据集与评估方式。本综述旨在为研究人员和实践者提供清晰的技术分类体系,并指明未来发展方向。更多论文列表与更新信息可访问 https://github.com/imxtx/awesome-controllabe-speech-synthesis。

原文摘要 · Abstract (English)

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has become a rapidly growing research area. This survey provides the first comprehensive review of controllable TTS methods, from traditional control techniques to emerging approaches using natural language prompts. We categorize model architectures, control strategies, and feature representations, while also summarizing challenges, datasets, and evaluations in controllable TTS. This survey aims to guide researchers and practitioners by offering a clear taxonomy and highlighting future directions in this fast-evolving field. One can visit https://github.com/imxtx/awesome-controllabe-speech-synthesis for a comprehensive paper list and updates.

语音合成大模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。