30B参数统一模型,让语音与文本共用一个大脑
Unified Audio Intelligence Without Regressing on Text Intelligence
- 用同一Transformer解码器处理音视频输入输出,实现无缝融合
- 1574亿音频+3205亿文本训练,性能全面领先且不损失文本能力
- 适合研究多模态生成、语音理解的开发者和工程师
音频智能涉及对音频和语音的理解、推理与生成。本文提出Nemotron-Labs-Audex-30B-A3B(Audex),基于强文本模型Nemotron-Cascade-2-30B-A3B构建的统一音文大模型。Audex采用统一架构:音频经编码后投影至文本嵌入空间,生成时音频与文本标记统一处理。该设计实现强大的音文融合、无缝多模态生成,并兼容标准LLM训练与推理基础设施。训练使用精心筛选的音文数据集,包含157.4B音频标记和320.5B文本标记,经过多阶段监督训练、文本仅有的Cascade RL及跨领域在线蒸馏。Audex在音频理解、语音识别与翻译、文本到语音、音频生成、语音到语音生成等任务上达到当前最优水平,同时保持其文本模型骨干的推理、对齐、知识、长上下文及代理能力,几乎无退化。模型检查点已公开,以促进开放研究。
原文摘要 · Abstract (English)
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a single Transformer decoder: audio inputs are encoded and projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. This architecture enables strong audio-text fusion, seamless multimodal generation, and compatibility with standard LLM training and inference infrastructure. For training, we meticulously curate audio-text datasets comprising 157.4B audio tokens and 320.5B text tokens. We apply multi-stage supervised training on these datasets, followed by text-only Cascade RL and multi-domain on-policy distillation. Audex delivers state-of-the-art audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, while preserving very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM backbone with marginal or no regression. We release the model checkpoints to facilitate open research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。