arXiv:2412.16677cs.CV2024-12被引 9

通过分阶段生成故事板与视频,实现可控且连贯的高质量视频生成。

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation

  • 先将文本转为包含人体姿态和物体布局的故事板,再据此生成视频。
  • 在VBench基准上,视觉质量和语义表达均优于现有方法。
  • 适合需要精准控制动作与场景结构的研究者或创作者。

从文本描述生成高质量视频面临保持时间连贯性和对主体运动控制的挑战。我们提出VAST(Video As Storyboard from Text),一个两阶段框架来解决这些问题,实现高质量视频生成。第一阶段,StoryForge将文本描述转化为详细的故事板,捕捉人体姿态和物体布局,以表征场景的结构本质。第二阶段,VisionForge基于这些故事板生成视频,产出具有流畅运动、时间一致性和空间一致性的高质量视频。通过解耦文本理解与视频生成,VAST实现了对主体动态和场景构图的精确控制。在VBench基准上的实验表明,VAST在视觉质量与语义表达方面均优于现有方法,为动态连贯视频生成树立了新标准。

原文摘要 · Abstract (English)

Generating high-quality videos from textual descriptions poses challenges in maintaining temporal coherence and control over subject motion. We propose VAST (Video As Storyboard from Text), a two-stage framework to address these challenges and enable high-quality video generation. In the first stage, StoryForge transforms textual descriptions into detailed storyboards, capturing human poses and object layouts to represent the structural essence of the scene. In the second stage, VisionForge generates videos from these storyboards, producing high-quality videos with smooth motion, temporal consistency, and spatial coherence. By decoupling text understanding from video generation, VAST enables precise control over subject dynamics and scene composition. Experiments on the VBench benchmark demonstrate that VAST outperforms existing methods in both visual quality and semantic expression, setting a new standard for dynamic and coherent video generation.

视频生成可控生成故事板

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。