用多模态信号统一生成音频,支持文本、视频、音频任意输入。
AudioX: A Unified Framework for Anything-to-Audio Generation
- 设计自适应融合模块,有效整合文本/视频/音频多模态输入
- 构建700万样本高质量数据集IF-caps,支持多任务训练
- 在文本到音频/音乐生成上超越现有方法,指令跟随能力强
基于灵活的多模态控制信号进行音频与音乐生成是一个广泛应用的课题,其核心挑战在于:1)统一的多模态建模范式;2)大规模、高质量的训练数据。为此,本文提出AudioX,一个面向任意模态到音频生成的统一框架,可整合文本、视频和音频等多种条件输入。该框架的核心是多模态自适应融合模块(Multimodal Adaptive Fusion),能有效融合多样化的多模态输入,增强跨模态对齐性并提升生成质量。为训练此统一模型,我们构建了一个大规模、高质量的数据集IF-caps,包含超过700万条样本,通过结构化标注流程精心采集,为多模态条件音频生成提供全面监督。我们在多种任务上对AudioX进行了基准测试,结果表明该模型在文本到音频及文本到音乐生成任务中均达到领先性能,充分展示了其在多模态控制信号下进行音频生成的能力,并展现出强大的指令遵循潜力。代码与数据集将公开于https://zeyuet.github.io/AudioX/。
原文摘要 · Abstract (English)
Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a unified multimodal modeling framework, and 2) large-scale, high-quality training data. As such, we propose AudioX, a unified framework for anything-to-audio generation that integrates varied multimodal conditions (i.e., text, video, and audio signals) in this work. The core design in this framework is a Multimodal Adaptive Fusion module, which enables the effective fusion of diverse multimodal inputs, enhancing cross-modal alignment and improving overall generation quality. To train this unified model, we construct a large-scale, high-quality dataset, IF-caps, comprising over 7 million samples curated through a structured data annotation pipeline. This dataset provides comprehensive supervision for multimodal-conditioned audio generation. We benchmark AudioX against state-of-the-art methods across a wide range of tasks, finding that our model achieves superior performance, especially in text-to-audio and text-to-music generation. These results demonstrate our method is capable of audio generation under multimodal control signals, showing powerful instruction-following potential. The code and datasets will be available at https://zeyuet.github.io/AudioX/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。