用自回归模型模拟摄影师构图意图,生成更自然的镜头运动轨迹。
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
- 基于导演视角设计的自回归模型,结合文本与深度图像生成镜头路径。
- 在29,000个真实镜头数据上训练,实现更精细、更稳定的镜头运动。
- 适合影视制作、AI视频生成领域研究者与创作者使用。
镜头轨迹设计在视频制作中至关重要,是传达导演意图和增强视觉叙事的核心工具。传统方法依赖几何优化或手工程序系统,而现有学习方法常存在结构偏差或缺乏文本对齐,限制了创造性表达。本文提出受摄影师启发的自回归模型GenDoP,首先构建包含29,000个真实镜头的多模态数据集DataDoP,涵盖自由移动镜头轨迹、深度图及详细描述动作、场景互动与导演意图的标注。基于此数据,训练一个仅解码器的Transformer模型,实现文本引导与RGBD输入下的高质量、上下文感知镜头运动生成。实验表明,相比现有方法,GenDoP具备更强可控性、更细粒度调整能力与更高运动稳定性。本工作为基于学习的电影摄影树立新标准,推动镜头控制与影视创作的未来发展。
原文摘要 · Abstract (English)
Camera trajectory design plays a crucial role in video production, serving as a fundamental tool for conveying directorial intent and enhancing visual storytelling. In cinematography, Directors of Photography meticulously craft camera movements to achieve expressive and intentional framing. However, existing methods for camera trajectory generation remain limited: Traditional approaches rely on geometric optimization or handcrafted procedural systems, while recent learning-based methods often inherit structural biases or lack textual alignment, constraining creative synthesis. In this work, we introduce an auto-regressive model inspired by the expertise of Directors of Photography to generate artistic and expressive camera trajectories. We first introduce DataDoP, a large-scale multi-modal dataset containing 29K real-world shots with free-moving camera trajectories, depth maps, and detailed captions in specific movements, interaction with the scene, and directorial intent. Thanks to the comprehensive and diverse database, we further train an auto-regressive, decoder-only Transformer for high-quality, context-aware camera movement generation based on text guidance and RGBD inputs, named GenDoP. Extensive experiments demonstrate that compared to existing methods, GenDoP offers better controllability, finer-grained trajectory adjustments, and higher motion stability. We believe our approach establishes a new standard for learning-based cinematography, paving the way for future advancements in camera control and filmmaking. Our project website: https://kszpxxzmc.github.io/GenDoP/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。