arXiv:2605.23463eess.AS2026-05被引 6

一个模型搞定语音识别、合成与实时对话,性能全超专用系统。

StepAudio 2.5 Technical Report

论文配图:StepAudio 2.5 Technical Report
图 1 · 摘自论文原文
  • 用人类反馈强化学习统一训练,按任务定制优化目标
  • 语音识别快而准,语音合成可控且有表现力,对话低延迟且人格一致
  • 适合想用单一模型解决多类语音任务的研发团队

统一音频-语言建模已成为现代语音系统的重要趋势,有望将大语言模型的推理能力引入听觉任务。然而,现有统一基础模型在自动语音识别(ASR)、文本转语音合成(TTS)和实时口语交互方面仍难以媲美专用系统,这一差距仍是开放挑战。本报告介绍 StepAudio 2.5,一个在三项能力上均达到或超越专用系统的统一音频-语言基础模型。我们不将这些任务视为架构差异,而是认为一旦文本与音频共享多模态表示空间,任务专化即为操作范式之别:数据构建、优化目标与解码约束。基于此洞察,我们将后训练范式从标准监督学习推进至任务定制化的基于人类反馈的强化学习(RLHF),以此作为定义复杂优化目标的核心机制。结合以RLHF为中心的对齐与专用解码策略,我们将共享主干塑造成三种不同运行模式:ASR分支通过可验证的多标记解码提升转录效率;TTS分支通过偏好驱动的RLHF与上下文丰富的监督实现可控且富有表现力的合成;实时分支则在RLHF框架内采用生成式奖励建模,实现低延迟、人格一致的对话。在标准基准测试中,StepAudio 2.5 在 ASR、TTS 和实时交互任务上均取得当前最佳性能,证明单一音频-语言基础模型可成功内化语音理解、生成与实时互动的不同部署目标。

原文摘要 · Abstract (English)

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks. However, existing unified foundations often struggle to match the depth of specialized systems across automatic speech recognition (ASR), text-to-speech synthesis (TTS), and realtime spoken interaction. Bridging this gap remains an open challenge. This report presents StepAudio 2.5, a unified audio-language foundation model that matches or exceeds specialized systems across all three capabilities. Rather than treating these tasks as architecturally distinct, we operate on the premise that once text and audio share a multimodal representational space, task specialization becomes a matter of operational regimes: data construction, optimization targets, and decoding constraints. Guided by this insight, we advance the post-training paradigm from standard supervised learning to task-tailored Reinforcement Learning from Human Feedback (RLHF), using it as the primary mechanism to define complex optimization targets. We leverage this RLHF-centric alignment, alongside specialized decoding, to shape a shared backbone into three distinct operational modes. Concretely, the ASR branch advances transcription efficiency via verifiable multi-token decoding; the TTS branch achieves controllable, expressive synthesis through preference-based RLHF and context-rich supervision; and the Realtime branch realizes low-latency, persona-consistent dialogue via generative reward modeling within an RLHF framework. On standard benchmarks, StepAudio 2.5 achieves state-of-the-art results across ASR, TTS, and Realtime, demonstrating that a singular audio-language foundation can successfully internalize the distinct deployment objectives of speech understanding, generation, and live interaction.

语音合成多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。