系统梳理语音对话模型发展,解析端到端技术与未来方向
WavChat: A Survey of Spoken Dialogue Models
- 按时间线分类语音对话系统,区分级联与端到端架构
- 总结语音表示、训练范式、实时交互等核心技术进展
- 适合研究者与开发者了解语音对话前沿与评测基准
近期以GPT-4o为代表的语音对话模型在语音领域引发广泛关注。相比传统三阶段级联架构(语音识别、大语言模型、文本转语音),现代语音对话模型展现出更强智能性,不仅能理解音频、音乐等语音特征,还可捕捉语音的风格与音色特点。同时支持低延迟、多轮高质量语音回复,实现听说同步的实时交互。尽管进展显著,现有研究仍缺乏对相关系统与技术的系统性综述。本文按时间顺序整理现有语音对话系统,将其划分为级联与端到端两类,并深入分析语音表征、训练范式、流式处理、双工交互等核心技术,指出各技术局限并提出未来研究方向。此外,全面回顾了训练与评估相关的数据集、评价指标与基准测试。相关材料可在https://github.com/jishengpeng/WavChat获取。
原文摘要 · Abstract (English)
Recent advancements in spoken dialogue models, exemplified by systems like GPT-4o, have captured significant attention in the speech domain. Compared to traditional three-tier cascaded spoken dialogue models that comprise speech recognition (ASR), large language models (LLMs), and text-to-speech (TTS), modern spoken dialogue models exhibit greater intelligence. These advanced spoken dialogue models not only comprehend audio, music, and other speech-related features, but also capture stylistic and timbral characteristics in speech. Moreover, they generate high-quality, multi-turn speech responses with low latency, enabling real-time interaction through simultaneous listening and speaking capability. Despite the progress in spoken dialogue systems, there is a lack of comprehensive surveys that systematically organize and analyze these systems and the underlying technologies. To address this, we have first compiled existing spoken dialogue systems in the chronological order and categorized them into the cascaded and end-to-end paradigms. We then provide an in-depth overview of the core technologies in spoken dialogue models, covering aspects such as speech representation, training paradigm, streaming, duplex, and interaction capabilities. Each section discusses the limitations of these technologies and outlines considerations for future research. Additionally, we present a thorough review of relevant datasets, evaluation metrics, and benchmarks from the perspectives of training and evaluating spoken dialogue systems. We hope this survey will contribute to advancing both academic research and industrial applications in the field of spoken dialogue systems. The related material is available at https://github.com/jishengpeng/WavChat.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。