arXiv:2504.06861cs.CVcs.AI2025-04CVPR被引 4

不改模型结构,用扩散轨迹交点生成视频,兼顾画面连贯与多样性。

EIDT-V: Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video Generation

  • 利用扩散轨迹交点与网格化提示控制帧间切换时机。
  • 在多个图像模型上实现最佳时序一致性与视觉质量。
  • 无需训练、适配广泛模型,适合快速部署文本到视频生成。

零样本、免训练、基于图像的文本到视频生成是新兴方向,旨在利用现有图像扩散模型生成视频。当前方法需对图像生成模型进行特定架构修改,限制了其适应性与可扩展性。本文提出一种模型无关的方法,仅基于潜空间中的扩散轨迹交点进行操作。单纯使用轨迹交点难以实现逐帧连贯性与多样性,因此采用网格化策略:通过上下文训练的LLM生成连贯的逐帧提示,并识别帧间差异;基于此构建CLIP注意力掩码,控制每个网格单元的提示切换时间。较早切换提升多样性,较晚切换增强连贯性。该方法在保持高灵活性的同时,实现当前最优的时序一致性、视觉保真度与用户满意度,为免训练、基于图像的文本到视频生成提供了新范式。

原文摘要 · Abstract (English)

Zero-shot, training-free, image-based text-to-video generation is an emerging area that aims to generate videos using existing image-based diffusion models. Current methods in this space require specific architectural changes to image generation models, which limit their adaptability and scalability. In contrast to such methods, we provide a model-agnostic approach. We use intersections in diffusion trajectories, working only with the latent values. We could not obtain localized frame-wise coherence and diversity using only the intersection of trajectories. Thus, we instead use a grid-based approach. An in-context trained LLM is used to generate coherent frame-wise prompts; another is used to identify differences between frames. Based on these, we obtain a CLIP-based attention mask that controls the timing of switching the prompts for each grid cell. Earlier switching results in higher variance, while later switching results in more coherence. Therefore, our approach can ensure appropriate control between coherence and variance for the frames. Our approach results in state-of-the-art performance while being more flexible when working with diverse image-generation models. The empirical analysis using quantitative metrics and user studies confirms our model's superior temporal consistency, visual fidelity and user satisfaction, thus providing a novel way to obtain training-free, image-based text-to-video generation.

文本生成视频扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。