用视觉块表示法实现复杂文本到视频的精准生成
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations

- 将视频分解为可控制的视觉块,实现对象级精细调控
- 在多个基准上达到顶尖布局可控性,零样本生成效果优秀
- 适合需要精确物体位置与运动控制的研究者与开发者
现有视频生成模型难以准确响应复杂文本提示并合成多个对象,亟需额外的定位输入以提升可控性。本文提出将视频分解为视觉原语——块视频表示(blob video representation),一种通用的可控视频生成表示。基于此,构建了名为BlobGEN-Vid的块引导视频扩散模型,支持用户对物体运动和细粒度外观进行控制。特别地,引入掩码3D注意力模块,有效提升帧间区域一致性;设计可学习插值模块,用于调节特定帧的文本嵌入,实现平滑的对象过渡。该框架具备模型无关性,基于U-Net与DiT两种视频扩散模型实现。大量实验表明,BlobGEN-Vid在多个基准上实现卓越的零样本视频生成能力与领先布局可控性。结合大语言模型进行布局规划时,其组合方案甚至超越专有文本到视频生成器的组合准确性。
原文摘要 · Abstract (English)
Existing video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompose videos into visual primitives - blob video representation, a general representation for controllable video generation. Based on blob conditions, we develop a blob-grounded video diffusion model named BlobGEN-Vid that allows users to control object motions and fine-grained object appearance. In particular, we introduce a masked 3D attention module that effectively improves regional consistency across frames. In addition, we introduce a learnable module to interpolate text embeddings so that users can control semantics in specific frames and obtain smooth object transitions. We show that our framework is model-agnostic and build BlobGEN-Vid based on both U-Net and DiT-based video diffusion models. Extensive experimental results show that BlobGEN-Vid achieves superior zero-shot video generation ability and state-of-the-art layout controllability on multiple benchmarks. When combined with an LLM for layout planning, our framework even outperforms proprietary text-to-video generators in terms of compositional accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。