用分镜稿生成连贯电影级多镜头叙事视频
STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative
- 以分镜结构替代稀疏关键帧,提升控制精度
- 跨镜头一致性与电影转场效果显著提升
- 适合影视创作与可控视频生成研究者
尽管生成模型在视频合成中已实现出色的视觉保真度,但生成连贯的多镜头叙事仍面临挑战。现有基于关键帧的方法虽具高效性与细粒度控制优势,却难以保持跨镜头一致性并捕捉电影语言。本文提出STAGE(Storyboard-Anchored GEneration)框架,将关键帧生成任务重构为分镜结构预测。我们设计了包含起止帧对的结构化分镜(STEP2),引入多镜头记忆包保障长程实体一致性,采用双编码策略增强单镜头内连贯性,并通过两阶段训练学习电影级镜头间过渡。我们还构建了大规模ConStoryBoard数据集,包含高质量电影片段及精细标注,涵盖故事进展、电影属性与人类偏好。大量实验表明,STAGE在结构化叙事控制与跨镜头一致性方面均表现更优。代码将公开。
原文摘要 · Abstract (English)
While recent advancements in generative models have achieved remarkable visual fidelity in video synthesis, creating coherent multi-shot narratives remains a significant challenge. To address this, keyframe-based approaches have emerged as a promising alternative to computationally intensive end-to-end methods, offering the advantages of fine-grained control and greater efficiency. However, these methods often fail to maintain cross-shot consistency and capture cinematic language. In this paper, we introduce STAGE, a SToryboard-Anchored GEneration workflow to reformulate the keyframe-based multi-shot video generation task. Instead of using sparse keyframes, we propose STEP2 to predict a structural storyboard composed of start-end frame pairs for each shot. We introduce the multi-shot memory pack to ensure long-range entity consistency, the dual-encoding strategy for intra-shot coherence, and the two-stage training scheme to learn cinematic inter-shot transition. We also contribute the large-scale ConStoryBoard dataset, including high-quality movie clips with fine-grained annotations for story progression, cinematic attributes, and human preferences. Extensive experiments demonstrate that STAGE achieves superior performance in structured narrative control and cross-shot coherence. Our code will be available at this url.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。