让视频生成模型像导演一样精准控制镜头,聚焦人物动作。
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation

- 用人类姿态定义镜头位置,实现以人物为中心的镜头设计
- 通过语言指令生成可预测的连续镜头轨迹,比现有方法更稳定
- 适合影视创作、动画设计等需要精准镜头控制的场景
生成式视频模型在视觉保真度和时间连贯性上已取得显著进展,但对镜头运动的主动控制仍难以实现。现有框架将镜头运动视为像素生成的副产物,导致轨迹随机、空间不一致且忽略以人物为核心的场景驱动。本文提出 Auteur,一种基于语言指令、以人物为中心的视频镜头生成方法。核心思想是:专业导演并非按世界坐标规划镜头,而是根据演员姿态定义画面构图,包括景别、角度和构图。我们将其形式化为以人体姿态为基准的镜头参数化方法,并引入可转换为标准6-自由度相机参数的领域专用语言(DSL)。一个微调后的多模态大语言模型充当虚拟导演,将自然语言描述与粗略的人体动作映射为稀疏的DSL关键帧,再通过确定性插值生成连续镜头轨迹,输入至视频生成器。我们在包含34,000组对齐数据的新数据集上训练与评估Auteur,数据源自程序合成及CondensedMovies真实电影片段。Auteur实现了以人为中心的电影级镜头控制,这是此前生成模型所缺乏的能力。我们提出新的聚焦于构图的评估指标,实验表明Auteur在各项指标上均优于现有方法。
原文摘要 · Abstract (English)
Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spatially inconsistent, and indifferent to the human subject driving the scene. In this work, we present Auteur, a method for language-driven, human-centric camera framing in generative video. Our core insight is that professional filmmakers conceive shots not as world-space trajectories but as framings defined relative to the actor, encoding shot size, angle, and composition as functions of human pose and motion. We formalize this intuition as a human-centric camera parameterization and introduce a Domain-Specific Language (DSL) that is convertible to standard 6-DoF camera parameters. A fine-tuned multimodal large language model then acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are deterministically interpolated into continuous camera trajectories, which are then provided as input to video generators. We train and evaluate Auteur on a new dataset of 34K aligned text, human motion, and DSL-annotated camera trajectories drawn from procedural synthesis and real-world movie footage from the CondensedMovies dataset. Auteur enables cinematographic framing of human-centered scenes, a capability largely absent in prior generative models. To assess this behavior, we propose new framing-focused metrics, and our experiments show that Auteur consistently outperforms existing methods. Project page is https://cyberiada.github.io/Auteur/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。