arXiv:2510.04577cs.SDcs.LG2025-10EMNLP被引 7

用反因果对齐让语言模型更好生成音频,性能超越扩散模型。

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

  • 设计多分支独立变压器,通过反因果对齐提升生成能力。
  • 在FMA和VCTK数据集上超越现有语言模型与扩散模型方法。
  • 适合追求高效文本到音频生成的开发者与研究者。

尽管语言模型(LM)结合残差向量量化(RVQ)分词器在文本到音频(T2A)生成中展现出潜力,但仍与基于扩散模型的方法存在显著差距。我们发现这一差距的核心矛盾在于:增加RVQ层可提升音频重建保真度,但超出传统语言模型的生成容量。分析表明:1)不同RVQ层间特征正交性阻碍语言模型训练;2)深层RVQ层令牌语义丰富度下降加剧自回归解码中的暴露偏差。为此,我们提出Siren框架,采用多个孤立的变压器,结合因果条件与强化学习实现的反因果对齐。大量实验表明,Siren在性能上超越现有的语言模型及扩散模型系统,达到当前最优水平。该方法融合语言模型的表征优势与音频合成的高保真需求,使语言模型在文本到音频任务中具备与扩散模型竞争的能力。同时,通过将音频表示与语言结构对齐,Siren为统一多模态生成框架提供了可行路径。

原文摘要 · Abstract (English)

While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin. We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improves audio reconstruction fidelity but exceeds the generation capacity of conventional LMs. To address this, we first analyze RVQ dynamics and uncover two key limitations: 1) orthogonality of features across RVQ layers hinders effective LMs training, and 2) descending semantic richness in tokens from deeper RVQ layers exacerbates exposure bias during autoregressive decoding. Based on these insights, we propose Siren, a novel LM-based framework that employs multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. Extensive experiments demonstrate that Siren outperforms both existing LM-based and diffusion-based T2A systems, achieving state-of-the-art results. By bridging the representational strengths of LMs with the fidelity demands of audio synthesis, our approach repositions LMs as competitive contenders against diffusion models in T2A tasks. Moreover, by aligning audio representations with linguistic structures, Siren facilitates a promising pathway toward unified multi-modal generation frameworks.

文本到音频语言模型反因果对齐生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。