FullDiT用全注意力机制统一多条件视频生成,解决控制冲突与冗余问题。
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
- 通过全注意力机制融合多任务条件,构建统一序列表示。
- 在多个数据集上达到当前最佳性能,生成质量优于适配器方法。
- 适合需要精细控制的视频生成研究者和工业应用开发者。
现有视频生成基础模型主要聚焦文本到视频任务,对细粒度内容创作控制有限。尽管基于适配器的方法(如ControlNet)可通过少量微调实现额外控制,但在集成多种条件时面临挑战:独立训练适配器间的分支冲突、参数冗余导致计算开销增加,且性能低于全微调。为此,我们提出FullDiT,一种通过统一全注意力机制无缝整合多条件的视频生成统一基础模型。通过将多任务条件融合为统一序列表示,并利用全自注意力的长程建模能力捕捉条件动态变化,FullDiT减少参数开销,避免条件冲突,展现可扩展性与涌现能力。我们进一步提出FullBench用于多任务视频生成评估。实验表明,FullDiT在多个基准上达到最先进水平,验证了全注意力在复杂多任务视频生成中的有效性。
原文摘要 · Abstract (English)
Current video generative foundation models primarily focus on text-to-video tasks, providing limited control for fine-grained video content creation. Although adapter-based approaches (e.g., ControlNet) enable additional controls with minimal fine-tuning, they encounter challenges when integrating multiple conditions, including: branch conflicts between independently trained adapters, parameter redundancy leading to increased computational cost, and suboptimal performance compared to full fine-tuning. To address these challenges, we introduce FullDiT, a unified foundation model for video generation that seamlessly integrates multiple conditions via unified full-attention mechanisms. By fusing multi-task conditions into a unified sequence representation and leveraging the long-context learning ability of full self-attention to capture condition dynamics, FullDiT reduces parameter overhead, avoids conditions conflict, and shows scalability and emergent ability. We further introduce FullBench for multi-task video generation evaluation. Experiments demonstrate that FullDiT achieves state-of-the-art results, highlighting the efficacy of full-attention in complex multi-task video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。