arXiv:2511.12921cs.CV2025-11被引 2

让AI视频能像专业摄影一样精准控制景深、快门等镜头效果。

Generative Photographic Control for Scene-Consistent Video Cinematic Editing

  • 用解耦注意力机制分离镜头运动与摄影参数,实现独立控制。
  • 在自建大规模数据集上训练,生成高质量且场景一致的视频。
  • 适合影视创作者、AI视频设计师快速实现电影级视觉风格。

电影叙事深受景深、曝光等摄影元素的艺术化操控影响,这些效果对传达情绪和营造美学至关重要。然而,现有生成视频模型大多仅支持相机运动控制,难以精细调节摄影参数。本文提出CineCtrl,首个支持专业相机参数(如虚化程度、快门速度)精细控制的视频电影编辑框架。通过引入解耦交叉注意力机制,将镜头运动与摄影输入分离,实现细粒度独立控制,同时保持场景一致性。为解决训练数据不足问题,构建了结合模拟摄影效果与真实采集的数据生成策略,形成大规模训练数据集。大量实验表明,该模型可生成高保真视频,精确实现用户指定的摄影效果。

原文摘要 · Abstract (English)

Cinematic storytelling is profoundly shaped by the artful manipulation of photographic elements such as depth of field and exposure. These effects are crucial in conveying mood and creating aesthetic appeal. However, controlling these effects in generative video models remains highly challenging, as most existing methods are restricted to camera motion control. In this paper, we propose CineCtrl, the first video cinematic editing framework that provides fine control over professional camera parameters (e.g., bokeh, shutter speed). We introduce a decoupled cross-attention mechanism to disentangle camera motion from photographic inputs, allowing fine-grained, independent control without compromising scene consistency. To overcome the shortage of training data, we develop a comprehensive data generation strategy that leverages simulated photographic effects with a dedicated real-world collection pipeline, enabling the construction of a large-scale dataset for robust model training. Extensive experiments demonstrate that our model generates high-fidelity videos with precisely controlled, user-specified photographic camera effects.

视频生成电影特效可控生成摄影控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。