arXiv:2412.02259cs.CVcs.AI2024-12被引 38

用一句话自动生成连贯多镜头视频,减少人工干预。

VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention

  • 分步构建故事脚本,自动规划角色、场景、镜头等要素。
  • 保持人物形象一致,支持表情和年龄变化的合理演进。
  • 无缝衔接镜头过渡,适合电影级长视频生成。

当前视频生成模型擅长短片段,但难以生成连贯的多镜头叙事,存在视觉断层与剧情断裂问题。现有方法或依赖大量人工编写,或牺牲跨镜头连续性以保证单镜头质量,限制了其在影视内容中的应用。本文提出VideoGen-of-Thought(VGoT),一种无需训练的分步框架,仅凭一句提示词即可自动生成多镜头视频。针对三大挑战:(1) 叙事碎片化,提出动态故事线建模,将用户输入转化为包含角色动态、背景延续、关系演变、镜头运动与HDR光照五个维度的详细分镜规范,并通过自验证确保逻辑连贯;(2) 视觉不一致,设计身份感知跨镜头传播机制,生成保留身份特征的人物令牌(IPP),支持表情、年龄等可控变化;(3) 转场缺陷,引入邻近潜在空间过渡机制,在镜头边界处处理特征重置,实现平滑视觉流。该系统在无需训练下,于面部一致性上超越基线20.4%,风格一致性提升17.4%,人工调整减少10倍,显著缩小了视觉合成与导演级叙事之间的差距。

原文摘要 · Abstract (English)

Current video generation models excel at short clips but fail to produce cohesive multi-shot narratives due to disjointed visual dynamics and fractured storylines. Existing solutions either rely on extensive manual scripting/editing or prioritize single-shot fidelity over cross-scene continuity, limiting their practicality for movie-like content. We introduce VideoGen-of-Thought (VGoT), a step-by-step framework that automates multi-shot video synthesis from a single sentence by systematically addressing three core challenges: (1) Narrative fragmentation: Existing methods lack structured storytelling. We propose dynamic storyline modeling, which turns the user prompt into concise shot drafts and then expands them into detailed specifications across five domains (character dynamics, background continuity, relationship evolution, camera movements, and HDR lighting) with self-validation to ensure logical progress. (2) Visual inconsistency: previous approaches struggle to maintain consistent appearance across shots. Our identity-aware cross-shot propagation builds identity-preserving portrait (IPP) tokens that keep character identity while allowing controlled trait changes (expressions, aging) required by the story. (3) Transition artifacts: Abrupt shot changes disrupt immersion. Our adjacent latent transition mechanisms implement boundary-aware reset strategies that process adjacent shots' features at transition points, enabling seamless visual flow while preserving narrative continuity. Combined in a training-free pipeline, VGoT surpasses strong baselines by 20.4\% in within-shot face consistency and 17.4\% in style consistency, while requiring 10x fewer manual adjustments. VGoT bridges the gap between raw visual synthesis and director-level storytelling for automated multi-shot video generation.

视频生成多镜头叙事连贯无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。