arXiv:2510.12000cs.SDcs.CL2025-10被引 16

一个模型搞定听觉理解、语音生成和跨模态推理,打破任务分割

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

  • 统一建模音频理解、文本转音频生成与多模态推理
  • 单模型在三大任务上达到顶尖专业模型水平
  • 首次实现音频领域的跨模态生成式推理,适合多任务研究者

当前音频语言建模研究将音频理解与文本转音频生成视为独立任务。极少工作尝试统一这些任务——而这正是迈向高级多模态推理的关键一步。本文提出统一音频语言模型(UALM),旨在单一模型中统一音频理解、文本转音频生成及多模态推理。我们首先提出UALM-Gen,一种直接预测音频标记的文本转音频语言模型,性能可媲美最先进的基于扩散的模型。通过合理的数据混合、训练策略和推理技术,我们的单一UALM模型在音频理解、文本转音频生成及文本推理三项任务上均达到当前最优专业模型的水平。此外,我们提出UALM-Reason,一种利用文本与音频共同进行中间思维推演的多模态推理模型,以支持复杂生成任务。据我们所知,这是音频研究中首次展示跨模态生成式推理,其有效性经主观评估验证。

原文摘要 · Abstract (English)

Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced multimodal reasoning. This paper introduces U}nified Audio Language Model (UALM), which aims to unify audio understanding, text-to-audio generation, and multimodal reasoning in a single model. To achieve this goal, we first present UALM-Gen, a text-to-audio language model that directly predicts audio tokens and is comparable to state-of-the-art diffusion-based models. We then demonstrate, using proper data blending, training recipes, and inference techniques, that our single UALM model matches the quality of state-of-the-art specialized models in audio understanding, text-to-audio generation, and text reasoning. Furthermore, we present UALM-Reason, a multimodal reasoning model that utilizes both text and audio in the intermediate thinking steps to facilitate complex generation tasks. To our knowledge, this is the first demonstration in audio research of cross-modal generative reasoning, with its effectiveness confirmed by subjective evaluations.

音频生成多模态语言模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。