一键生成连贯多镜头视频,自动解决角色一致性与转场问题。
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
- 分步构建故事线,动态生成镜头细节
- 角色形象跨镜头保持一致,支持表情变化
- 自适应转场机制,减少90%以上人工修改
当前视频生成模型擅长短片段,但在多镜头叙事上因视觉不连贯和剧情断裂而表现不佳。现有方法或依赖大量人工脚本,或牺牲跨镜头连续性以保证单镜头质量。本文提出VideoGen-of-Thought(VGoT),通过三步机制实现从单句指令自动生成多镜头视频:(1)叙事碎片化:引入动态故事线建模,将用户输入转化为包含角色动态、背景延续、关系演变、镜头运动与HDR光照等五个维度的详细分镜描述,确保逻辑连贯并支持自我验证;(2)视觉不一致:提出身份感知跨镜头传播,生成保持角色特征的面孔令牌(IPP),在遵循剧情的前提下允许表情、年龄等变化;(3)过渡伪影:设计邻接潜在空间过渡机制,在切换点处进行边界感知重置,实现视觉流畅且叙事连贯的转场。实验表明,VGoT在单镜头人脸一致性上优于基线20.4%,风格一致性提升17.4%,跨镜头一致性提高超100%,人工调整次数仅为其他方法的1/10。
原文摘要 · Abstract (English)
Current video generation models excel at short clips but fail to produce cohesive multi-shot narratives due to disjointed visual dynamics and fractured storylines. Existing solutions either rely on extensive manual scripting/editing or prioritize single-shot fidelity over cross-scene continuity, limiting their practicality for movie-like content. We introduce VideoGen-of-Thought (VGoT), a step-by-step framework that automates multi-shot video synthesis from a single sentence by systematically addressing three core challenges: (1) Narrative Fragmentation: Existing methods lack structured storytelling. We propose dynamic storyline modeling, which first converts the user prompt into concise shot descriptions, then elaborates them into detailed, cinematic specifications across five domains (character dynamics, background continuity, relationship evolution, camera movements, HDR lighting), ensuring logical narrative progression with self-validation. (2) Visual Inconsistency: Existing approaches struggle with maintaining visual consistency across shots. Our identity-aware cross-shot propagation generates identity-preserving portrait (IPP) tokens that maintain character fidelity while allowing trait variations (expressions, aging) dictated by the storyline. (3) Transition Artifacts: Abrupt shot changes disrupt immersion. Our adjacent latent transition mechanisms implement boundary-aware reset strategies that process adjacent shots' features at transition points, enabling seamless visual flow while preserving narrative continuity. VGoT generates multi-shot videos that outperform state-of-the-art baselines by 20.4% in within-shot face consistency and 17.4% in style consistency, while achieving over 100% better cross-shot consistency and 10x fewer manual adjustments than alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。