构建语音分词器评估基准,揭示其性能差异与机制特点。
STAB: Speech Tokenizer Assessment Benchmark
- 提出STAB框架,统一评估多种语音分词器的特性。
- 在5个下游任务中验证,不同分词器表现差异显著。
- 适合研究语音-文本对齐与大模型适配的研究者使用。
将语音表示为离散符号,可将其转换为类似文本的格式,从而让语音能作为大型语言模型(LLMs)的输入。尽管已有多种语音分词器被提出,但针对特定下游任务所需的分词器属性及其整体泛化能力仍不明确。跨多个下游任务评估分词器性能计算开销大,难以扩展。为此,我们提出STAB(Speech Tokenizer Assessment Benchmark),一个系统化的评估框架,用于全面评估语音分词器并揭示其内在特性。该框架深化了对语音分词机制的理解,为未来分词器模型的发展提供宝贵资源,并支持基于标准化基准的对比分析。我们在多种语音任务和分词器选择下评估了STAB指标,并将其与下游任务性能相关联。
原文摘要 · Abstract (English)
Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (LLMs). Currently, while several speech tokenizers have been proposed, there is ambiguity regarding the properties that are desired from a tokenizer for specific downstream tasks and its overall generalizability. Evaluating the performance of tokenizers across different downstream tasks is a computationally intensive effort that poses challenges for scalability. To circumvent this requirement, we present STAB (Speech Tokenizer Assessment Benchmark), a systematic evaluation framework designed to assess speech tokenizers comprehensively and shed light on their inherent characteristics. This framework provides a deeper understanding of the underlying mechanisms of speech tokenization, thereby offering a valuable resource for expediting the advancement of future tokenizer models and enabling comparative analysis using a standardized benchmark. We evaluate the STAB metrics and correlate this with downstream task performance across a range of speech tasks and tokenizer choices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。