arXiv:2502.06490eess.AScs.AI2025-02TPAMI综述被引 81

梳理离散语音标记最新进展,揭示其在语音生成中的核心作用

Recent Advances in Discrete Speech Tokens: A Review

  • 按声学与语义特性分类,系统归纳离散语音标记方法
  • 对比不同标记类型在语音建模中的性能表现与适用场景
  • 指出当前挑战并提出未来研究方向,适合语音模型开发者参考

大型语言模型(LLMs)时代下,语音生成技术飞速发展,离散语音标记已成为语音表示的基础范式。这类标记具有离散、紧凑、简洁的特点,不仅利于高效传输与存储,且天然兼容语言建模范式,使语音可无缝融入以文本为主的LLM架构。现有研究将离散语音标记分为声学标记与语义标记两大类,各自形成丰富研究领域,具备独特的设计哲学与方法路径。本文系统综述了离散语音标记的现有分类体系与近期创新,批判性分析各类范式的优劣,并开展跨标记类型的系统实验比较。此外,识别出该领域的持续挑战并提出潜在研究方向,旨在为离散语音标记的未来发展与应用提供切实可行的洞见。

原文摘要 · Abstract (English)

The rapid advancement of speech generation technologies in the era of large language models (LLMs) has established discrete speech tokens as a foundational paradigm for speech representation. These tokens, characterized by their discrete, compact, and concise nature, are not only advantageous for efficient transmission and storage, but also inherently compatible with the language modeling framework, enabling seamless integration of speech into text-dominated LLM architectures. Current research categorizes discrete speech tokens into two principal classes: acoustic tokens and semantic tokens, each of which has evolved into a rich research domain characterized by unique design philosophies and methodological approaches. This survey systematically synthesizes the existing taxonomy and recent innovations in discrete speech tokenization, conducts a critical examination of the strengths and limitations of each paradigm, and presents systematic experimental comparisons across token types. Furthermore, we identify persistent challenges in the field and propose potential research directions, aiming to offer actionable insights to inspire future advancements in the development and application of discrete speech tokens.

语音生成离散标记大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。