arXiv:2506.10274cs.SDcs.AI2025-06综述被引 64

系统评测离散音频标记器,助力语音与语言模型融合。

Discrete Audio Tokens: More Than a Survey!

  • 构建跨语音、音乐、通用音频的标记化方法分类体系
  • 在重建、下游任务与声学建模上评估多个标记器性能
  • 揭示关键局限与未来方向,适合音频-语言模型研究者

离散音频标记是紧凑的表示形式,旨在保留感知质量、语音内容和说话人特征的同时,实现高效存储与推理,并在多样下游任务中表现良好。它们为连续特征提供了实用替代方案,使语音与音频可融入现代大语言模型(LLMs)。随着基于标记的音频处理兴起,各类标记化方法涌现,已有若干综述涵盖最新进展。但现有研究多聚焦特定领域或任务,缺乏跨基准的统一比较。本文提出对离散音频标记器的系统性综述与基准测试,覆盖语音、音乐与通用音频三个领域。我们基于编码器-解码器结构、量化技术、训练范式、流式能力及应用领域构建分类体系,评估标记器在重建、下游性能与声学语言建模等多基准上的表现,并通过受控消融实验分析权衡关系。研究揭示关键局限、实用考量与开放挑战,为该快速发展的领域提供洞见与指导。更多详情(含主要结果与标记器数据库)请访问:https://poonehmousavi.github.io/dates-website/

原文摘要 · Abstract (English)

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse downstream tasks. They provide a practical alternative to continuous features, enabling the integration of speech and audio into modern large language models (LLMs). As interest in token-based audio processing grows, various tokenization methods have emerged, and several surveys have reviewed the latest progress in the field. However, existing studies often focus on specific domains or tasks and lack a unified comparison across various benchmarks. This paper presents a systematic review and benchmark of discrete audio tokenizers, covering three domains: speech, music, and general audio. We propose a taxonomy of tokenization approaches based on encoder-decoder, quantization techniques, training paradigm, streamability, and application domains. We evaluate tokenizers on multiple benchmarks for reconstruction, downstream performance, and acoustic language modeling, and analyze trade-offs through controlled ablation studies. Our findings highlight key limitations, practical considerations, and open challenges, providing insight and guidance for future research in this rapidly evolving area. For more information, including our main results and tokenizer database, please refer to our website: https://poonehmousavi.github.io/dates-website/.

音频标记大模型语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。