arXiv:2508.08244cs.CVcs.AI2025-08SIGGRAPH被引 17

用上下文调优生成符合电影剪辑逻辑的下一镜头。

Cut2Next: Generating Next Shot via In-Context Tuning

  • 通过分层提示词引导扩散Transformer生成下一镜头。
  • 在多个数据集上实现视觉一致性和文本匹配度提升。
  • 适合影视生成、剧本可视化等需要连贯叙事的场景。

高效多镜头生成需具备电影式转场与严格的影像连贯性。现有方法多关注基础视觉一致性,忽视关键剪辑模式(如正反打、切出镜头),导致输出虽视觉连贯却缺乏叙事深度与电影质感。为此,我们提出下一镜头生成(NSG):合成一个高质量且严格遵循专业剪辑模式的后续镜头。框架Cut2Next采用扩散变换器(DiT),结合新型分层多提示策略,通过关系提示定义整体语境与镜头间剪辑风格,个体提示指定每镜头内容与电影摄影属性,共同引导生成。架构创新包括上下文感知条件注入(CACI)与分层注意力掩码(HAM),在不增加参数的前提下融合多元信号。我们构建了原始版RawCuts与精炼版CuratedCuts数据集,并引入CutBench评估基准。实验表明,Cut2Next在视觉一致性和文本保真度上表现优异;用户研究显示其在剪辑模式遵循度与整体电影连贯性上显著更受青睐,验证其生成高质量、具叙事表达力与电影完整性的后续镜头的能力。

原文摘要 · Abstract (English)

Effective multi-shot generation demands purposeful, film-like transitions and strict cinematic continuity. Current methods, however, often prioritize basic visual consistency, neglecting crucial editing patterns (e.g., shot/reverse shot, cutaways) that drive narrative flow for compelling storytelling. This yields outputs that may be visually coherent but lack narrative sophistication and true cinematic integrity. To bridge this, we introduce Next Shot Generation (NSG): synthesizing a subsequent, high-quality shot that critically conforms to professional editing patterns while upholding rigorous cinematic continuity. Our framework, Cut2Next, leverages a Diffusion Transformer (DiT). It employs in-context tuning guided by a novel Hierarchical Multi-Prompting strategy. This strategy uses Relational Prompts to define overall context and inter-shot editing styles. Individual Prompts then specify per-shot content and cinematographic attributes. Together, these guide Cut2Next to generate cinematically appropriate next shots. Architectural innovations, Context-Aware Condition Injection (CACI) and Hierarchical Attention Mask (HAM), further integrate these diverse signals without introducing new parameters. We construct RawCuts (large-scale) and CuratedCuts (refined) datasets, both with hierarchical prompts, and introduce CutBench for evaluation. Experiments show Cut2Next excels in visual consistency and text fidelity. Crucially, user studies reveal a strong preference for Cut2Next, particularly for its adherence to intended editing patterns and overall cinematic continuity, validating its ability to generate high-quality, narratively expressive, and cinematically coherent subsequent shots.

视频生成扩散模型剪辑逻辑叙事连贯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。