arXiv:2605.23508cs.GRcs.AI2026-05被引 1

用草图分镜控制长视频生成,实现动作与画面的精准协同。

DrawVideo: Generating Long Video from Storyboard Keyframe Sketches

论文配图:DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
图 1 · 摘自论文原文
  • 以黑白草图+提示词分解长视频为可独立控制的镜头
  • 生成视频保持结构连贯、画面稳定且角色一致
  • 适合需要精细控制动作和场景的动画创作人员

长视频生成需高保真合成、连贯叙事结构及对长时间跨度的用户控制。现有文本到视频方法多依赖单一长提示词,限制了对姿态、构图、布局和运动的控制。我们提出 DrawVideo,一种草图引导、分镜驱动的可控长视频生成框架。DrawVideo 将长视频分解为独立可控的镜头,每个镜头由黑白草图、外观提示词和运动提示词定义:草图控制姿态与布局,外观提示词定义身份、场景与风格,运动提示词引导时间动态。采用分层‘全局多镜头、局部单草图’策略:先生成结构对齐的参考关键帧,再将运动提示扩展为表示动作状态的衍生关键帧,最后在相邻关键帧间合成片段完成每段镜头。我们还构建了首个草图引导的文本到长视频数据集 SketchLongVideo,通过动画视频的镜头检测、关键帧提取、视觉-语言识别、提示词分解与草图转换获得。实验表明,DrawVideo 实现了强结构可控性、外观一致性、视觉稳定性与连贯长视频生成。

原文摘要 · Abstract (English)

Long video generation requires high-fidelity synthesis, coherent narrative structure, and user control over extended time spans. Existing text-to-video methods often rely on a single long prompt, limiting control over pose, composition, layout, and motion. We propose DrawVideo, a sketch-guided, storyboard-driven framework for controllable long-video generation. DrawVideo decomposes long videos into independently controllable shots, each defined by a black-and-white sketch, an appearance prompt, and a motion prompt. The sketch controls pose and layout, the appearance prompt defines identity, scene, and style, and the motion prompt guides temporal dynamics. DrawVideo follows a hierarchical 'global multi-shot, local single-sketch' strategy: it first generates a structure-aligned reference keyframe, then expands the motion prompt into derivative keyframes representing action states, and finally synthesizes clips between adjacent keyframes to build each shot. We also introduce SketchLongVideo, the first dataset for sketch-guided text-to-long-video generation, constructed from animation videos via shot detection, keyframe extraction, vision-language recognition, prompt decomposition, and sketch conversion. Experiments show that DrawVideo achieves strong structural controllability, appearance consistency, visual stability, and coherent long-video generation.

长视频生成草图控制分镜驱动动画创作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。