用真实镜头几何信息生成连贯短剧,提升画面与剧情一致性。
DramaDirector: Geometry-Guided Short Drama Generation

- 通过深度与姿态索引真实镜头库,引导分镜构图与生成。
- 在35部剧集、8.1万镜头上验证,多指标优于现有方法。
- 适合影视生成、剧本可视化及视频创作研究者使用。
短剧以快速剪辑节奏、对话驱动的视角切换和高要求的影像表现著称,现有文本或逐级视频生成方案难以满足。本文研究从剧情到短剧的生成任务,将全局剧情与局部上下文转化为具视觉基础的多镜头视频。提出DramaDirector框架,利用真实短剧镜头库中按深度与姿态索引的影像几何信息,指导分镜生成。该框架将每帧分解为静态视觉与动态叙事条件,采用基于语义模板的SFT与GRPO训练,在学习到的图文对齐奖励下优化规划器,并检索深度-姿态参考以引导首帧生成与图像到视频合成。同时构建DramaBoard基准,涵盖35部真人短剧、2.8千集、8.1万镜头,提供结构化分镜与多维评估协议。实验表明,DramaDirector在忠实性、一致性和可控性上均优于代表性多智能体与视频生成基线。
原文摘要 · Abstract (English)
Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation pipelines struggle to meet. We study plot-to-short-drama generation, where a global plot and local context are transformed into visually grounded multi-shot videos. We propose DramaDirector, a geometry-grounded framework that lets the planner borrow cinematographic geometry from a gallery of real short-drama shots indexed by depth and pose. DramaDirector decouples each shot into static visual and dynamic narrative conditions, trains the planner with schema-constrained SFT and GRPO under a learned text-visual alignment reward, and retrieves depth-pose references to guide first-frame generation and image-to-video synthesis. We also introduce DramaBoard, a benchmark built from 35 live-action dramas, 2.8K episodes, and 81K shots, with structured storyboards and multi-dimensional evaluation protocols. Experiments show that DramaDirector improves over representative multi-agent and video generation baselines on faithfulness, consistency, and controllability. Our code is released at: https://github.com/iLearn-Lab/DramaDirector
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。