arXiv:2412.12223cs.CVcs.AI2024-12被引 5

让AI生成视频更有电影感,能自动拍出专业镜头语言。

Can video generation replace cinematographers? Research on the cinematic language of generated video

  • 构建20类电影语言数据集,教会AI理解镜头构图与运镜。
  • 提出CameraDiff和CLIPLoRA,实现精准可控的镜头切换与风格融合。
  • 适合影视创作、AI视频生成研究者,提升自动化视频表现力。

文本到视频(T2V)生成近年来借助扩散模型显著提升了视频的视觉连贯性,但现有研究多聚焦物体运动,忽视了对情感传达与叙事节奏至关重要的电影语言。为此,我们提出三重方法:首先,构建包含20个子类别的精细标注电影语言数据集,涵盖镜头构图、视角与摄像机运动,使模型可学习多样化的电影风格;其次,提出CameraDiff,采用LoRA实现精确且稳定的电影化控制,支持灵活镜头生成;第三,设计CameraCLIP以评估电影语义对齐并指导多镜头组合。基于CameraCLIP,进一步提出CLIPLoRA——一种受CLIP引导的动态LoRA组合方法,可自适应融合多个预训练电影化LoRA,实现镜头间平滑过渡与无缝风格混合。实验表明,CameraDiff保证了稳定精准的电影控制,CameraCLIP达到R@1为0.83,CLIPLoRA显著提升单视频内多镜头组合质量,缩小了自动化视频生成与专业电影制作之间的差距。

原文摘要 · Abstract (English)

Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing research primarily focuses on object motion, often overlooking cinematic language, which is crucial for conveying emotion and narrative pacing in cinematography. To address this, we propose a threefold approach to improve cinematic control in T2V models. First, we introduce a meticulously annotated cinematic language dataset with twenty subcategories, covering shot framing, shot angles, and camera movements, enabling models to learn diverse cinematic styles. Second, we present CameraDiff, which employs LoRA for precise and stable cinematic control, ensuring flexible shot generation. Third, we propose CameraCLIP, designed to evaluate cinematic alignment and guide multi-shot composition. Building on CameraCLIP, we introduce CLIPLoRA, a CLIP-guided dynamic LoRA composition method that adaptively fuses multiple pre-trained cinematic LoRAs, enabling smooth transitions and seamless style blending. Experimental results demonstrate that CameraDiff ensures stable and precise cinematic control, CameraCLIP achieves an R@1 score of 0.83, and CLIPLoRA significantly enhances multi-shot composition within a single video, bridging the gap between automated video generation and professional cinematography.\textsuperscript{1}

视频生成电影语言扩散模型LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。