用关键帧控制叙事节奏,生成更像电影的视频。
SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control

- 通过多个关键帧引导视频生成,实现精准叙事控制。
- 在多镜头合成与视频扩展任务中优于现有方法。
- 适合影视创作、智能编剧等需要节奏把控的场景。
视频的叙事质量从根本上决定了其感知价值。尽管现有视频生成方法能产出视觉吸引人的内容,但主要依赖文本提示或首/末帧等稀疏条件信号,难以精确控制叙事结构与时间节奏。本文提出 SmartDirector 框架,通过多个关键帧增强视频生成模型的叙事能力。该框架支持单镜头生成、多镜头叙事合成与视频扩展等多种灵活场景。系统分为两个阶段:Director-Gen 在关键帧条件下生成低分辨率视频;Director-SR 利用高分辨率关键帧作为语义锚点,恢复细节。为支持多关键帧训练,构建了从电影中提取单镜头与多镜头序列的数据管道。大量实验表明,SmartDirector 显著优于当前最先进方法。代码将公开以促进后续研究。
原文摘要 · Abstract (English)
The narrative quality of a video fundamentally determines its perceptual value. Although existing video generation methods can produce visually appealing content, they predominantly rely on sparse conditioning signals such as text prompts or first/last frames, which limits precise control over narrative structure and temporal pacing. In this paper, we propose SmartDirector, a framework that enhances the narrative capacity of video generation models through multiple keyframes. SmartDirector supports flexible generation scenarios including single-shot generation, multi-shot narrative synthesis, and video extension. The framework operates in two stages: Director-Gen generates a low-resolution video conditioned on the provided keyframes, and Director-SR refines the output by exploiting high-resolution keyframes as semantic anchors to recover fine-grained details. To enable robust multi-keyframe training, we construct a data pipeline that curates single-shot and multi-shot sequences from movies. Extensive experiments demonstrate that SmartDirector substantially outperforms existing state-of-the-art approaches. We will release the code to facilitate further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。