统一生成音视频的扩散模型,支持多任务、高多样性输出。
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
- 用共享隐空间和统一去噪网络,联合建模音视频相关性。
- 在三个任务上逼近顶尖单任务模型性能,生成内容更贴近真实数据分布。
- 适合需要多模态生成、多样化输出的研究与应用开发。
随着扩散模型的发展,音视频生成技术取得显著进步。然而,现有方法大多依赖各模态独立模块,缺乏统一生成架构;且多局限于单一任务和小规模数据集。为此,我们提出UniForm——一种统一的多任务扩散变换器,在共享隐空间中同时生成音频与视觉内容。通过统一去噪网络,模型捕捉音视频间的内在关联。我们还设计了任务特定噪声方案与任务标记,使模型仅用一组参数即可支持视频到音频、音频到视频及文本到音视频生成。结合大语言模型与大规模文本-音频-视频联合数据集,UniForm实现了比以往方法更高的生成多样性。实验表明,UniForm在三项生成任务上表现接近当前最优单任务模型,生成内容不仅高度符合真实数据分布,还能实现更丰富、更精细的生成。
原文摘要 · Abstract (English)
With the rise of diffusion models, audio-video generation has been revolutionized. However, most existing methods rely on separate modules for each modality, with limited exploration of unified generative architectures. In addition, many are confined to a single task and small-scale datasets. To overcome these limitations, we introduce UniForm, a unified multi-task diffusion transformer that generates both audio and visual modalities in a shared latent space. By using a unified denoising network, UniForm captures the inherent correlations between sound and vision. Additionally, we propose task-specific noise schemes and task tokens, enabling the model to support multiple tasks with a single set of parameters, including video-to-audio, audio-to-video and text-to-audio-video generation. Furthermore, by leveraging large language models and a large-scale text-audio-video combined dataset, UniForm achieves greater generative diversity than prior approaches. Experiments show that UniForm achieves performance close to the state-of-the-art single-task models across three generation tasks, with generated content that is not only highly aligned with real-world data distributions but also enables more diverse and fine-grained generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。