arXiv:2502.11946cs.CLcs.AI2025-02被引 120

开源语音交互模型Step-Audio实现统一理解与生成,支持动态控制和多模态任务。

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

  • 130B参数多模态模型统一处理语音与文本,支持对话与生成
  • 通过轻量级语音克隆生成3B模型,在指令跟随上提升9.3%准确率
  • 支持方言、情绪、说唱等动态调整,适合智能语音助手开发

实时语音交互是人机协作的基础接口,但现有开源模型存在语音数据成本高、动态控制弱、智能程度不足等问题。本文提出首个可投入生产的开源方案Step-Audio,包含:1)1300亿参数的统一语音-文本多模态模型,支持理解与生成,其Chat版本已开源;2)生成式语音数据引擎,构建低成本语音克隆框架,并通过蒸馏生成30亿参数的轻量级Step-Audio-TTS-3B模型;3)指令驱动的精细控制机制,支持方言、情绪、歌唱与说唱等动态调节;4)增强认知架构,集成工具调用与角色扮演能力,有效处理复杂任务。基于新构建的StepEval-Audio-360评估基准,Step-Audio在人类评测中达到领先性能,尤其在指令遵循方面表现突出。在公开基准LLaMA Question上平均提升9.3%。代码与模型已开源于https://github.com/stepfun-ai/Step-Audio。

原文摘要 · Abstract (English)

Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such as high costs in voice data collection, weakness in dynamic control, and limited intelligence. To address these challenges, this paper introduces Step-Audio, the first production-ready open-source solution. Key contributions include: 1) a 130B-parameter unified speech-text multi-modal model that achieves unified understanding and generation, with the Step-Audio-Chat version open-sourced; 2) a generative speech data engine that establishes an affordable voice cloning framework and produces the open-sourced lightweight Step-Audio-TTS-3B model through distillation; 3) an instruction-driven fine control system enabling dynamic adjustments across dialects, emotions, singing, and RAP; 4) an enhanced cognitive architecture augmented with tool calling and role-playing abilities to manage complex tasks effectively. Based on our new StepEval-Audio-360 evaluation benchmark, Step-Audio achieves state-of-the-art performance in human evaluations, especially in terms of instruction following. On open-source benchmarks like LLaMA Question, shows 9.3% average performance improvement, demonstrating our commitment to advancing the development of open-source multi-modal language technologies. Our code and models are available at https://github.com/stepfun-ai/Step-Audio.

语音生成多模态指令控制开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。