arXiv:2502.19548cs.CLcs.SD2025-02ACL综述被引 13

梳理大模型与语音融合的三种主流方法,揭示技术路径与应用前景。

When Large Language Models Meet Speech: A Survey on Integration Approaches

  • 按文本、隐式表征、音频令牌三类方式融合语音与大模型
  • 覆盖语音识别、对话系统等多场景应用,验证方法有效性
  • 适合关注多模态交互、语音智能的研究者参考

大语言模型(LLMs)的快速发展推动其从纯文本任务向多模态扩展。大量研究探索将语音等其他模态与LLMs融合,其中语音因与文本天然相关而成为重点。本文系统综述了语音与大模型的融合方法,将其分为三类:基于文本的方法、基于隐式表示的方法和基于音频令牌的方法。文章进一步展示了这些方法在语音识别、语音合成、语音对话等应用场景中的实践,并指出了当前面临的挑战,为后续研究提供方向与启发。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have spurred interest in expanding their application beyond text-based tasks. A large number of studies have explored integrating other modalities with LLMs, notably speech modality, which is naturally related to text. This paper surveys the integration of speech with LLMs, categorizing the methodologies into three primary approaches: text-based, latent-representation-based, and audio-token-based integration. We also demonstrate how these methods are applied across various speech-related applications and highlight the challenges in this field to offer inspiration for

大模型语音融合多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。