arXiv:2601.13910eess.AScs.SD2026-01综述被引 6

系统梳理深度学习歌声合成技术,助你快速掌握该领域全貌。

Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches

  • 按任务类型分类现有系统,归纳级联与端到端两大架构范式。
  • 深入分析歌声建模与控制核心技术,涵盖训练策略与评估标准。
  • 适合语音合成、音乐生成研究者及工程师快速入门参考。

近年来,歌声合成(SVS)的进展吸引了学术界和工业界的广泛关注。随着大语言模型和新型生成范式的出现,可控且高保真的歌声生成已成为可能。然而,该领域仍缺乏对基于深度学习的歌声合成系统及其支撑技术的系统性综述。为此,本文首先按任务类型对现有系统进行分类,并将当前架构归纳为级联与端到端两大范式。进一步,我们深入分析了核心关键技术,包括歌声建模与控制方法。最后,系统回顾了相关数据集、标注工具和评估基准,支持模型训练与性能评测。附录中还介绍了训练策略及更深入讨论。本综述提供了对最新SVS研究文献的全面回顾,对研究人员和工程师具有重要参考价值。相关资料可在 https://github.com/David-Pigeon/SyntheticSingers 获取。

原文摘要 · Abstract (English)

Recent advances in singing voice synthesis (SVS) have attracted substantial attention from both academia and industry. With the advent of large language models and novel generative paradigms, producing controllable, high-fidelity singing voices has become an attainable goal. Yet the field still lacks a comprehensive survey that systematically analyzes deep-learning-based singing voice synthesis systems and their enabling technologies. To address the aforementioned issue, this survey first categorizes existing systems by task type and then organizes current architectures into two major paradigms: cascaded and end-to-end approaches. Moreover, we provide an in-depth analysis of core technologies, covering singing modeling and control techniques. Finally, we review relevant datasets, annotation tools, and evaluation benchmarks that support training and assessment. In appendix, we introduce training strategies and further discussion of SVS. This survey provides an up-to-date review of the literature on SVS models, which would be a useful reference for both researchers and engineers. Related materials are available at https://github.com/David-Pigeon/SyntheticSingers.

歌声合成深度学习综述语音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。