arXiv:2410.18908eess.AS2024-10中稿 · publication as a S…综述被引 110

系统梳理语音大模型的架构与挑战,助力更懂人类语言的智能交互。

A Survey on Speech Large Language Models for Understanding

  • 提出语音理解三维度分类框架,明确任务目标
  • 分析语音大模型三阶段抽象架构,揭示技术演进路径
  • 指出指令敏感与语义推理退化问题,指引未来方向

语音理解对于解析口语中蕴含的语言、副语言及非语言信息至关重要,是实现高效人机交互的基础。大型语言模型(LLM)的快速发展催生了语音大语言模型(Speech LLMs),标志着通用语音理解系统的范式转变。本文正式定义语音理解概念,构建涵盖信息、功能与格式维度的结构化分类体系。在此框架下,全面综述当前Speech LLMs,从模态特征提取、模态信息融合到大模型推理三个阶段分析其架构设计;同时探讨训练策略、代表性数据集与评估方法。基于实证分析与实验结果,识别出两大核心挑战:指令敏感性及语义推理能力下降,并提出具体改进方向。本研究旨在为构建更鲁棒、可泛化且对齐人类认知的Speech LLMs提供基础参考。

原文摘要 · Abstract (English)

Speech understanding is essential for interpreting the diverse forms of information embedded in spoken language, including linguistic, paralinguistic, and non-linguistic cues that are vital for effective human-computer interaction. The rapid advancement of large language models (LLMs) has catalyzed the emergence of Speech Large Language Models (Speech LLMs), which marks a transformative shift toward general-purpose speech understanding systems. To further clarify and systematically delineate task objectives, in this paper, we formally define the concept of speech understanding and introduce a structured taxonomy encompassing its informational, functional, and format dimensions. Within this scope of definition, we present a comprehensive review of current Speech LLMs, analyzing their architectures through a three-stage abstraction: Modality Feature Extraction, Modality Information Fusion, and LLM Inference. In addition, we examine training strategies, discuss representative datasets, and review evaluation methodologies adopted in the field. Based on empirical analyses and experimental evidence, we identify two key challenges currently facing Speech LLMs: instruction sensitivity and degradation in semantic reasoning and propose concrete directions for addressing these issues. Through this systematic and detailed survey, we aim to offer a foundational reference for researchers and practitioners working toward more robust, generalizable, and human-aligned Speech LLMs.

语音理解大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。