构建语音大模型发展路线图,推动超人类语音理解。
Roadmap towards Superhuman Speech Understanding using Large Language Models
- 提出从语音识别到超人类理解的五级发展路径
- 发现语音大模型在语调线索和声学知识上存在能力缺口
- 设计SAGI基准测试评估多层级语音理解能力
大型语言模型(LLM)的成功推动了语音与音频数据的融合,旨在构建能处理文本与非文本输入的通用基础模型。近期进展如GPT-4o展示了端到端语音大模型的潜力,可保留非语义信息与世界知识以实现深层语音理解。为指导语音大模型的发展,本文提出一个五级路线图,涵盖从基础自动语音识别(ASR)到能够整合非语义信息与抽象声学知识以完成复杂任务的超人类模型。同时,我们设计了SAGI基准,统一标准化各层级关键任务的评估标准,揭示了在使用抽象声学知识和能力完整性方面的挑战。研究发现当前模型在处理副语言线索和抽象声学知识方面仍存明显差距,并提出了未来发展方向。本文系统梳理了语音大模型的发展路径,引入评估基准,并提供了对现有局限与潜力的关键洞见。
原文摘要 · Abstract (English)
The success of large language models (LLMs) has prompted efforts to integrate speech and audio data, aiming to create general foundation models capable of processing both textual and non-textual inputs. Recent advances, such as GPT-4o, highlight the potential for end-to-end speech LLMs, which preserves non-semantic information and world knowledge for deeper speech understanding. To guide the development of speech LLMs, we propose a five-level roadmap, ranging from basic automatic speech recognition (ASR) to advanced superhuman models capable of integrating non-semantic information with abstract acoustic knowledge for complex tasks. Moreover, we design a benchmark, SAGI Bechmark, that standardizes critical aspects across various tasks in these five levels, uncovering challenges in using abstract acoustic knowledge and completeness of capability. Our findings reveal gaps in handling paralinguistic cues and abstract acoustic knowledge, and we offer future directions. This paper outlines a roadmap for advancing speech LLMs, introduces a benchmark for evaluation, and provides key insights into their current limitations and potential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。