对比离散与连续语音特征在语义任务中的表现,发现连续特征更优。
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
- 用轻量级模型对比离散与连续语音特征在多种语义任务上的表现。
- 连续特征在细粒度语义理解任务中显著优于离散令牌。
- 揭示离散令牌性能不足源于粒度有限和信息保留效率低。
随着语音大语言模型(Speech LLMs)的兴起,离散语音标记因其能与文本标记无缝集成而受到关注。与多数聚焦于连续语音特征的研究不同,尽管基于离散标记的LLMs在某些任务上表现出色,但两种范式之间的性能差距尚未被充分探讨。本文使用轻量级模型(Qwen1.5-0.5B)在多种语义相关任务中进行了公平、全面的比较。结果表明,连续特征总体上优于离散标记,尤其在需要精细语义理解的任务中。此外,本研究深入分析了离散标记表现不佳的关键因素,如标记粒度受限和信息保留效率低下。基于此分析,我们探索了提升离散标记性能的潜在方向。希望本研究为推进语音大模型中离散语音标记的发展提供新视角。
原文摘要 · Abstract (English)
With the rise of Speech Large Language Models (Speech LLMs), there has been growing interest in discrete speech tokens for their ability to integrate with text-based tokens seamlessly. Compared to most studies that focus on continuous speech features, although discrete-token based LLMs have shown promising results on certain tasks, the performance gap between these two paradigms is rarely explored. In this paper, we present a fair and thorough comparison between discrete and continuous features across a variety of semantic-related tasks using a light-weight LLM (Qwen1.5-0.5B). Our findings reveal that continuous features generally outperform discrete tokens, particularly in tasks requiring fine-grained semantic understanding. Moreover, this study goes beyond surface-level comparison by identifying key factors behind the under-performance of discrete tokens, such as limited token granularity and inefficient information retention. To enhance the performance of discrete tokens, we explore potential aspects based on our analysis. We hope our results can offer new insights into the opportunities for advancing discrete speech tokens in Speech LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。