arXiv:2604.09195cs.AI2026-04被引 3

用多智能体模拟电影拍摄流程,生成连贯有表现力的叙事视频。

Camera Artist: A Multi-Agent Framework for Cinematic Language Storytelling Video Generation

论文配图:Camera Artist: A Multi-Agent Framework for Cinematic Language Storytelling Video Generation
图 1 · 摘自论文原文
  • 引入摄制镜头智能体,递归生成分镜增强镜头间叙事连贯性
  • 通过电影语言注入,提升镜头设计的表现力与电影质感
  • 适合影视生成、创意内容创作人群使用

我们提出 Camera Artist,一个模拟真实电影制作流程的多智能体框架,用于生成具有明确电影语言的叙事视频。尽管近期多智能体系统在从剧本到视频的自动化制作上取得进展,但普遍缺乏对相邻镜头间叙事推进的显式机制和对电影语言的刻意运用,导致故事断裂、影片质量有限。为解决此问题,Camera Artist 在现有智能体流水线基础上,引入专用的摄制镜头智能体,融合递归分镜生成以强化镜头间叙事连续性,并通过电影语言注入实现更具表现力、更贴近电影风格的镜头设计。大量定量与定性实验表明,该方法在叙事一致性、动态表现力和感知电影质量方面均显著优于现有基线。

原文摘要 · Abstract (English)

We propose Camera Artist, a multi-agent framework that models a real-world filmmaking workflow to generate narrative videos with explicit cinematic language. While recent multi-agent systems have made substantial progress in automating filmmaking workflows from scripts to videos, they often lack explicit mechanisms to structure narrative progression across adjacent shots and deliberate use of cinematic language, resulting in fragmented storytelling and limited filmic quality. To address this, Camera Artist builds upon established agentic pipelines and introduces a dedicated Cinematography Shot Agent, which integrates recursive storyboard generation to strengthen shot-to-shot narrative continuity and cinematic language injection to produce more expressive, film-oriented shot designs. Extensive quantitative and qualitative results demonstrate that our approach consistently outperforms existing baselines in narrative consistency, dynamic expressiveness, and perceived film quality.

视频生成多智能体电影语言叙事连贯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。