对比离散令牌与连续特征在语音大模型中的表现,发现连续特征更优。
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
- 统一实验条件下比较自监督学习的离散与连续语音表征
- 连续特征在六项语音理解任务中整体表现更优
- 适合关注语音建模策略优化的研究者参考
随着语音大语言模型(SpeechLLMs)的发展,语音处理主要分为离散令牌和连续特征两种范式。尽管两者在音频任务中均表现出色,但二者之间的性能差距尚未得到充分探讨。为此,我们在相同实验设置下,对基于自监督学习(SSL)的离散与连续特征进行公平比较。使用小规模(Qwen1.5-0.5B)和大规模(Llama3.1-8B)语言模型,在六个语音理解相关任务上评估其表现,并进一步开展高效性对比、SSL层分析、LLM层分析及鲁棒性测试。结果表明,连续特征在多数任务中表现更优。不同语音处理方法在信息学习与处理模式上呈现显著差异。研究结果有望为提升语音大模型的语音理解能力提供重要参考。
原文摘要 · Abstract (English)
With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features. Each approach has demonstrated strong capabilities in audio-related processing tasks. However, the performance gap between these two paradigms has not been thoroughly explored. To address this gap, we present a fair comparison of self-supervised learning (SSL)-based discrete and continuous features under the same experimental settings. We evaluate their performance across six spoken language understanding-related tasks using both small and large-scale LLMs (Qwen1.5-0.5B and Llama3.1-8B). We further conduct in-depth analyses, including efficient comparison, SSL layer analysis, LLM layer analysis, and robustness comparison. Our findings reveal that continuous features generally outperform discrete tokens in various tasks. Each speech processing method exhibits distinct characteristics and patterns in how it learns and processes speech information. We hope our results will provide valuable insights to advance spoken language understanding in SpeechLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。