arXiv:2604.09803cs.SD2026-04

一个框架搞定音乐生成与音源分离,支持多种输入条件。

MAGE: Modality-Agnostic Music Generation and Target-Source Extraction

  • 统一框架用同一模型处理文本、视觉或混合输入的音乐生成。
  • 在音源分离任务中,干扰抑制能力提升,有效提取目标音轨。
  • 适合需要多模态条件生成与提取的研究者使用。

近期多模态音频生成进展实现了从文本、视觉等高层条件生成音乐,但多数系统仅针对单一任务:无参考混合音轨时生成音乐,或从已有混合音轨中提取目标音源。这种固定任务设计限制了其在多种输入组合下的应用。为此,我们提出MAGE,一种在共享连续潜在空间内实现条件音乐生成与基于混合音轨的目标音源提取的模态无关框架。方法包括:1)可控多模态流变换器建模从噪声到目标音频潜在表示的条件流,使同一主干网络可在有或无混合条件时运行;2)音频-视觉关联对齐将帧级视觉特征映射至音频潜在序列,实现音频令牌级别的视觉条件调节;3)交叉门控调制机制利用对齐后的视觉表示调控中间音频特征,同时由文本提供独立语义引导。通过动态模态掩码训练,模型同时接触文本独有、视觉独有、文本-视觉联合、混合条件及无条件配置。在MUSIC基准上,分别评估了无混合生成和基于混合的音源提取任务。结果表明,MAGE在两种设置下均提供统一的条件接口,且所提对齐与门控组件显著提升了提取任务中的干扰抑制能力。

原文摘要 · Abstract (English)

Recent advances in multimodal audio generation have enabled music synthesis from text, visual cues, and other high-level conditions. However, most systems are designed for a single operating mode: either generating music without a reference mixture or extracting a target source from an existing mixture. This fixed-task design limits their use when different combinations of text, visual, and mixture inputs are available. To address this gap, we propose MAGE, a modality-agnostic framework for conditional music generation and mixture-grounded target-source extraction within a shared continuous latent space. Our approach introduces three key components. First, a Controlled Multimodal FluxFormer models the conditional flow from noise to a target audio latent, enabling the same backbone to operate with or without a mixture condition. Second, Audio-Visual Nexus Alignment maps frame-level visual features onto the audio latent sequence, allowing visual evidence to condition the generation process at the audio-token level. Third, a cross-gated modulation mechanism uses the aligned visual representation to regulate intermediate audio features, while text provides separate semantic guidance. We further train MAGE with dynamic modality masking, exposing the same model to text-only, visual-only, joint text-visual, mixture-conditioned, and unconditional configurations. Experiments on the MUSIC benchmark evaluate MAGE under separate protocols for mixture-free generation and mixture-grounded target-source extraction. The results show that MAGE provides a shared conditioning interface across both settings, and that the proposed alignment and gating components improve interference suppression in the extraction task.

音乐生成音源分离多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。