arXiv:2607.26529cs.CV2026-07被引 1

无需训练即可生成长时序、多镜头且可控的电影级视频。

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

论文配图:CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling
图 1 · 摘自论文原文
  • 通过调整位置编码和注意力模式,打破时间连续性以实现清晰转场。
  • 支持参考控制,每镜头可精细调节角色与场景,保持全局视觉一致性。
  • 适合影视创作、视频编辑等需要长视频可控生成的场景。

电影视频生成对文本到视频扩散模型构成挑战,因其需同时满足多镜头生成、角色与场景的细粒度控制以及跨长时程的连贯生成。现有方法依赖定制化与重训练分别应对特定需求,无法在统一框架下兼顾全部要求。本文提出无需训练的CineWeaver框架,核心洞察是预训练视频扩散模型存在时间连续性结构偏差导致多镜头生成困难。通过操纵推理阶段的位置编码与注意力模式,打破时间连续性以实现清晰转场;进一步引入按镜头路由的参考条件机制实现每镜头细粒度控制,并设计锚点记忆机制保障长视频生成中的全局外观一致性。据我们所知,CineWeaver是首个在无需训练的前提下,统一实现长时序、参考可控、多镜头视频生成的框架。实验表明,该方法可生成高质量、身份一致、全局外观稳定且转场清晰的长视频。项目页:https://cineweaver.github.io。

原文摘要 · Abstract (English)

Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately address specific requirements, and cannot simultaneously fulfill all the requirements with a unified framework. In this paper, we shed light on the training-free paradigm with the key insight that the difficulty of multi-shot generation arises from a structural bias toward temporal continuity in pretrained video diffusion models, and consequently, propose a unified framework named CineWeaver to achieve reference-controllable multi-shot long-video generation without retraining. We manipulate positional encoding and attention patterns to break temporal continuity during inference to enable clear shot transitions using pretrained video diffusion models. Furthermore, we extend the proposed framework with a shot-routed reference conditioning mechanism for per-shot fine-grained controllability, and develop an anchor memory mechanism to allow long-form generation with consistent global appearance cues. To our best knowledge, CineWeaver is the first unified framework to simultaneously enable \textbf{long-form}, \textbf{reference-controllable}, and \textbf{multi-shot} video generation in a training-free fashion. Experimental results demonstrate that CineWeaver produces high-quality cinematic videos of long durations with consistent identities, stable global appearance, and clear shot transitions. The project page is available at: https://cineweaver.github.io.

视频生成参考控制长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。