arXiv:2505.22053cs.SDcs.MA2025-05被引 20

无需训练的多智能体系统,实现图文视频生成多样音频。

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation

  • 分层多智能体架构,动态分配专家模型处理不同音频类型。
  • 通过试错迭代优化,提升音频生成质量与上下文对齐度。
  • 首个针对多模态到多音频生成的基准数据集,支持9项指标评估。

多模态输入(如视频、文本、图像)生成多样化音频(如音效、语音、音乐、歌曲)面临挑战,主要源于高质量配对数据稀缺及缺乏稳健的多任务学习框架。近期多智能体系统展现出潜力,但直接应用于多模态到多音频生成存在三大问题:(1)对多模态输入(尤其是视频)理解不够细粒度;(2)单模型难以应对多样音频事件;(3)缺乏输出自纠错机制。为此,我们提出 AudioGenie,一种无需训练的多智能体框架,采用双层架构——生成团队与监督团队。生成团队包含细粒度任务分解与自适应 Mixture-of-Experts(MoE)协作机制,实现全面多模态理解与动态模型选择,并引入试错式迭代优化模块实现自我修正。监督团队通过反馈回路确保时空一致性并验证输出。此外,我们构建了 MA-Bench,首个面向多模态到多音频生成的基准,包含198个标注视频及多类型音频。实验表明,AudioGenie 在8个任务中9项指标达到或超过当前最优(SOTA)水平。用户研究进一步验证了方法在质量、准确性、对齐度与审美上的有效性。项目主页与音频样例见 https://audiogenie.github.io/。

原文摘要 · Abstract (English)

Multimodality-to-Multiaudio (MM2MA) generation faces significant challenges in synthesizing diverse and contextually aligned audio types (e.g., sound effects, speech, music, and songs) from multimodal inputs (e.g., video, text, images), owing to the scarcity of high-quality paired datasets and the lack of robust multi-task learning frameworks. Recently, multi-agent system shows great potential in tackling the above issues. However, directly applying it to MM2MA task presents three critical challenges: (1) inadequate fine-grained understanding of multimodal inputs (especially for video), (2) the inability of single models to handle diverse audio events, and (3) the absence of self-correction mechanisms for reliable outputs. To this end, we propose AudioGenie, a novel training-free multi-agent system featuring a dual-layer architecture with a generation team and a supervisor team. For the generation team, a fine-grained task decomposition and an adaptive Mixture-of-Experts (MoE) collaborative entity are designed for detailed comprehensive multimodal understanding and dynamic model selection, and a trial-and-error iterative refinement module is designed for self-correction. The supervisor team ensures temporal-spatial consistency and verifies outputs through feedback loops. Moreover, we build MA-Bench, the first benchmark for MM2MA tasks, comprising 198 annotated videos with multi-type audios. Experiments demonstrate that our AudioGenie achieves state-of-the-art (SOTA) or comparable performance across 9 metrics in 8 tasks. User study further validates the effectiveness of our method in terms of quality, accuracy, alignment, and aesthetic. The project website with audio samples can be found at https://audiogenie.github.io/.

音频生成多智能体多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。