arXiv:2604.10708cs.SDcs.AI2026-04International Conf…被引 9

首个统一音频生成与编辑的端到端框架,支持通用声音、音乐和语音。

Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing

论文配图:Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
图 1 · 摘自论文原文
  • 融合冻结多模态大模型与可训练扩散变换器,实现跨域统一建模。
  • 在百万级音频编辑数据集上达到领先性能,超越专用模型。
  • 支持零样本跨语言控制,适合通用音频智能研究者使用。

多模态模型的进步推动了音频理解、生成与编辑的快速发展。然而,这些能力通常由专用模型处理,真正能无缝整合三者的统一框架仍待探索。尽管已有工作尝试统一音频理解与生成,但大多局限于特定领域。为此,我们提出 Audio-Omni,首个端到端框架,实现通用声音、音乐和语音领域的生成与编辑统一,并集成多模态理解能力。其架构结合冻结的多模态大语言模型进行高层推理,以及可训练的扩散变换器实现高保真合成。为解决音频编辑中数据稀缺问题,我们构建了 AudioEdit,一个包含超过一百万条精心标注编辑对的大规模数据集。大量实验表明,Audio-Omni在多项基准测试中达到顶尖水平,优于以往统一方法,性能媲美甚至超越专用专家模型。此外,Audio-Omni展现出知识增强推理生成、上下文生成及零样本跨语言控制等继承能力,揭示了通用生成式音频智能的可行方向。代码、模型与数据集将公开发布于 https://zeyuet.github.io/Audio-Omni。

原文摘要 · Abstract (English)

Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized models, leaving the development of a truly unified framework that can seamlessly integrate all three tasks underexplored. While some pioneering works have explored unifying audio understanding and generation, they often remain confined to specific domains. To address this, we introduce Audio-Omni, the first end-to-end framework to unify generation and editing across general sound, music, and speech domains, with integrated multi-modal understanding capabilities. Our architecture synergizes a frozen Multimodal Large Language Model for high-level reasoning with a trainable Diffusion Transformer for high-fidelity synthesis. To overcome the critical data scarcity in audio editing, we construct AudioEdit, a new large-scale dataset comprising over one million meticulously curated editing pairs. Extensive experiments demonstrate that Audio-Omni achieves state-of-the-art performance across a suite of benchmarks, outperforming prior unified approaches while achieving performance on par with or superior to specialized expert models. Beyond its core capabilities, Audio-Omni exhibits remarkable inherited capabilities, including knowledge-augmented reasoning generation, in-context generation, and zero-shot cross-lingual control for audio generation, highlighting a promising direction toward universal generative audio intelligence. The code, model, and dataset will be publicly released on https://zeyuet.github.io/Audio-Omni.

音频生成多模态扩散模型统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。