arXiv:2604.17195cs.CV2026-04中稿 · CVPR

用视频扩散模型生成连贯分镜,人物一致且叙事流畅。

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

论文配图:DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior
图 1 · 摘自论文原文
  • 基于视频扩散先验,支持文本或参考图生成分镜。
  • 多参考图身份对齐,角色一致性显著提升。
  • 适合影视创作、动画设计等需要连贯视觉叙事的场景。

分镜合成在视觉叙事中至关重要,旨在生成连贯的镜头序列,以一致的角色、场景和转场视觉化电影情节。然而,现有方法大多基于文本到图像扩散模型,难以维持长时序连贯性、角色身份一致性和跨镜头叙事流。本文提出 DreamShot,一种基于视频生成模型的分镜框架,充分挖掘视频扩散先验,实现可控的多镜头合成。DreamShot 支持文本到镜头(Text-to-Shot)与参考图到镜头(Reference-to-Shot)生成,以及基于前序帧的故事续写,实现灵活且上下文感知的分镜生成。通过利用视频生成模型固有的时空一致性,DreamShot 生成了在视觉与语义上更连贯的序列,提升了叙事保真度与角色连续性。此外,其引入多参考角色条件模块,接受多个角色参考图像,并通过角色注意力一致性损失强制参考与生成角色间的注意力对齐,显式约束角色身份。大量实验表明,相比最先进文本到图像分镜模型,DreamShot 在场景连贯性、角色一致性与生成效率上均表现更优,为可控视频模型驱动的视觉叙事开辟新方向。

原文摘要 · Abstract (English)

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly adapted from text-to-image diffusion models, which struggle to maintain long-range temporal coherence, consistent character identities, and narrative flow across multiple shots. In this paper, we introduce DreamShot, a video generative model based storyboard framework that fully exploits powerful video diffusion priors for controllable multi-shot synthesis. DreamShot supports both Text-to-Shot and Reference-to-Shot generation, as well as story continuation conditioned on previous frames, enabling flexible and context-aware storyboard generation. By leveraging the spatial-temporal consistency inherent in video generative models, DreamShot produces visually and semantically coherent sequences with improved narrative fidelity and character continuity. Furthermore, DreamShot incorporates a multi-reference role conditioning module that accepts multiple character reference images and enforces identity alignment via a Role-Attention Consistency Loss, explicitly constraining attention between reference and generated roles. Extensive experiments demonstrate that DreamShot achieves superior scene coherence, role consistency, and generation efficiency compared to state-of-the-art text-to-image storyboard models, establishing a new direction toward controllable video model-driven visual storytelling.

视频生成分镜合成扩散模型角色一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。