端到端音频理解模型,能听懂语气情绪并实时对话
Step-Audio 2 Technical Report
- 用隐变量音频编码+强化学习,让模型理解说话风格和情绪
- 生成离散音频令牌,响应速度更快,支持变声切换
- 结合外部工具检索与搜索,减少幻觉,适合真实场景对话
本文介绍 Step-Audio 2,一个面向工业级音频理解与语音对话的端到端多模态大语言模型。通过整合隐变量音频编码器与以推理为核心的强化学习,该模型在自动语音识别(ASR)和音频理解任务中表现优异。为实现真正的端到端语音对话,Step-Audio 2将离散音频令牌生成融入语言建模,显著提升对语调、语气等副语言信息的响应能力。为充分利用真实数据中的丰富文本与声学知识,模型引入检索增强生成(RAG),可调用网络搜索等外部工具以降低幻觉风险,并通过音频搜索实现音色切换。模型基于数百万小时的真实语音与音频数据训练,在多种音频理解与对话评测中超越其他开源及商用方案,达到当前最佳水平。
原文摘要 · Abstract (English)
This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent audio encoder and reasoning-centric reinforcement learning (RL), Step-Audio 2 achieves promising performance in automatic speech recognition (ASR) and audio understanding. To facilitate genuine end-to-end speech conversation, Step-Audio 2 incorporates the generation of discrete audio tokens into language modeling, significantly enhancing its responsiveness to paralinguistic information such as speaking styles and emotions. To effectively leverage the rich textual and acoustic knowledge in real-world data, Step-Audio 2 integrates retrieval-augmented generation (RAG) and is able to call external tools such as web search to mitigate hallucination and audio search to switch timbres. Trained on millions of hours of speech and audio data, Step-Audio 2 delivers intelligence and expressiveness across diverse conversational scenarios. Evaluation results demonstrate that Step-Audio 2 achieves state-of-the-art performance on various audio understanding and conversational benchmarks compared to other open-source and commercial solutions. Please visit https://github.com/stepfun-ai/Step-Audio2 for more information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。