arXiv:2602.16687cs.SDcs.CL2026-02被引 3

构建能同时理解语音语义、声学细节和文本的音频大模型,支持通用音频生成与跨模态任务。

Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens

  • 采用多模态令牌混合策略,联合建模语义、声学与文本信息。
  • 在64个模型上验证缩放规律,发现最优数据量增长速度是模型规模的1.6倍。
  • 训练出从135M到4B参数的SODA系列模型,可直接用于语音翻译等任务。

当前音频语言模型多以文本为主导,或仅依赖语义音频令牌,限制了通用音频建模能力。本文系统研究了原生音频基础模型在大规模下应用下一步令牌预测的方法,联合建模语义内容、声学细节与文本,以支持通用音频生成和跨模态能力。我们提供三项关键发现:(1) 系统评估了数据来源、文本混合比例与令牌构成等设计选择,确立了可复现的训练方案;(2) 首次通过同FLOP分析对64个模型(涵盖$3{ imes}10^{18}$至$3{ imes}10^{20}$ FLOPs)开展缩放律研究,发现最优数据量增长速度比最优模型规模快1.6倍;(3) 基于上述规律训练出SODA(Scaling Open Discrete Audio)系列模型,参数规模为135M至4B,使用500B令牌进行训练,其性能符合预测结果且优于现有模型。该模型可作为统一架构灵活适配多种音频/文本任务,例如通过微调实现保留音色的语音到语音翻译。

原文摘要 · Abstract (English)

Current audio language models are predominantly text-first, either extending pre-trained text LLM backbones or relying on semantic-only audio tokens, limiting general audio modeling. This paper presents a systematic empirical study of native audio foundation models that apply next-token prediction to audio at scale, jointly modeling semantic content, acoustic details, and text to support both general audio generation and cross-modal capabilities. We provide comprehensive empirical insights for building such models: (1) We systematically investigate design choices -- data sources, text mixture ratios, and token composition -- establishing a validated training recipe. (2) We conduct the first scaling law study for discrete audio models via IsoFLOP analysis on 64 models spanning $3{\times}10^{18}$ to $3{\times}10^{20}$ FLOPs, finding that optimal data grows 1.6$\times$ faster than optimal model size. (3) We apply these lessons to train SODA (Scaling Open Discrete Audio), a suite of models from 135M to 4B parameters on 500B tokens, comparing against our scaling predictions and existing models. SODA serves as a flexible backbone for diverse audio/text tasks -- we demonstrate this by fine-tuning for voice-preserving speech-to-speech translation, using the same unified architecture.

音频生成基础模型跨模态语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。