arXiv:2504.08528cs.CLcs.SD2025-04综述被引 151

系统梳理语音语言模型发展脉络与关键技术

On The Landscape of Spoken Language Models: A Comprehensive Survey

  • 按架构、训练、评估三维度统一分类近年研究
  • 揭示语音模型从专用到通用的演进趋势
  • 适合关注语音AI系统化发展的研究人员

语音语言处理正从定制化、任务专用模型转向通用语音语言模型(SLMs),类似文本自然语言处理中通用语言模型的发展。SLMs包括纯语音语言模型(建模语音分词序列分布)以及结合语音编码器与文本语言模型的混合架构,支持语音与文本的输入输出。该领域研究多样,术语和评估方式不统一。本文通过系统综述,梳理近年进展,按模型架构、训练策略与评估方法对工作进行分类,并探讨关键挑战与未来方向。

原文摘要 · Abstract (English)

The field of spoken language processing is undergoing a shift from training custom-built, task-specific models toward using and optimizing spoken language models (SLMs) which act as universal speech processing systems. This trend is similar to the progression toward universal language models that has taken place in the field of (text) natural language processing. SLMs include both "pure" language models of speech -- models of the distribution of tokenized speech sequences -- and models that combine speech encoders with text language models, often including both spoken and written input or output. Work in this area is very diverse, with a range of terminology and evaluation settings. This paper aims to contribute an improved understanding of SLMs via a unifying literature survey of recent work in the context of the evolution of the field. Our survey categorizes the work in this area by model architecture, training, and evaluation choices, and describes some key challenges and directions for future work.

语音模型语言模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。