arXiv:2501.12909cs.CLcs.GR2025-01被引 19

用多个AI角色协作自动拍虚拟电影,效果比单个模型更好。

FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D Spaces

  • 构建导演、编剧、演员等角色的多智能体系统,协同完成从创意到拍摄全流程。
  • 在15个创意上生成视频,人类评分平均3.98分(满分5分),优于所有基线。
  • 即使使用较弱的GPT-4o模型,仍超越单智能体o1,展现协同优势。

虚拟电影制作涉及复杂的决策过程,包括剧本创作、虚拟摄影和演员位置与动作的精准设定。受语言智能体社会自动化决策进展的启发,本文提出FilmAgent——一种基于大语言模型的多智能体协作框架,实现我们构建的3D虚拟空间中的端到端电影自动化。FilmAgent模拟导演、编剧、演员、摄影师等多种剧组角色,覆盖电影制作的关键阶段:(1) 创意开发将头脑风暴转化为结构化故事大纲;(2) 剧本创作细化每场戏的对白与角色动作;(3) 摄影确定每镜头的摄像机设置。一组智能体通过迭代反馈与修订协作,验证中间脚本并减少幻觉。我们在15个创意上评估生成视频,涵盖4个关键方面。人工评价显示,FilmAgent在所有方面均优于所有基线,平均得分3.98(满分5分),证明了多智能体协作在影视制作中的可行性。进一步分析表明,尽管使用较弱的GPT-4o模型,FilmAgent仍超越单智能体o1,凸显了良好协调的多智能体系统的优势。最后,我们讨论了OpenAI文本生成视频模型Sora与FilmAgent在影视制作中的互补性优缺点。

原文摘要 · Abstract (English)

Virtual film production requires intricate decision-making processes, including scriptwriting, virtual cinematography, and precise actor positioning and actions. Motivated by recent advances in automated decision-making with language agent-based societies, this paper introduces FilmAgent, a novel LLM-based multi-agent collaborative framework for end-to-end film automation in our constructed 3D virtual spaces. FilmAgent simulates various crew roles, including directors, screenwriters, actors, and cinematographers, and covers key stages of a film production workflow: (1) idea development transforms brainstormed ideas into structured story outlines; (2) scriptwriting elaborates on dialogue and character actions for each scene; (3) cinematography determines the camera setups for each shot. A team of agents collaborates through iterative feedback and revisions, thereby verifying intermediate scripts and reducing hallucinations. We evaluate the generated videos on 15 ideas and 4 key aspects. Human evaluation shows that FilmAgent outperforms all baselines across all aspects and scores 3.98 out of 5 on average, showing the feasibility of multi-agent collaboration in filmmaking. Further analysis reveals that FilmAgent, despite using the less advanced GPT-4o model, surpasses the single-agent o1, showing the advantage of a well-coordinated multi-agent system. Lastly, we discuss the complementary strengths and weaknesses of OpenAI's text-to-video model Sora and our FilmAgent in filmmaking.

多智能体虚拟制作内容生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。