arXiv:2604.25318cs.GRcs.AI2026-04被引 1

用大模型自动生成游戏过场动画,实现编剧到特效全流程闭环控制。

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation

论文配图:Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation
图 1 · 摘自论文原文
  • 构建双向交互框架,让大模型能实时观察并操控游戏引擎中的场景状态。
  • 多智能体协同工作,导演统筹动画、摄影、音效,视觉反馈持续优化效果。
  • 推出新评测基准,专门测试长时序、多步骤的过场动画生成能力。

过场动画是视频游戏与互动媒体中精心编排的电影式序列,承担叙事传递、角色塑造与情感共鸣的核心功能。制作过程高度复杂,需剧本、摄影、角色动画、配音与技术指导等多领域协作,通常需数天至数周团队努力才能产出数分钟高质量内容。本文提出Cutscene Agent,一个用于自动化端到端过场动画生成的大模型智能体框架。该框架有三项贡献:(1) 基于模型上下文协议(MCP)的过场工具包,实现大模型智能体与游戏引擎之间的双向联动——智能体不仅能调用引擎操作,还能持续观测实时场景状态,支持可编辑的原生引擎资产闭环生成;(2) 多智能体系统,由导演智能体协调动画、摄影、音效等专业子智能体,并引入视觉推理反馈回路,实现基于感知的精细化优化;(3) CutsceneBench,一个分层评估基准,用于衡量过场动画生成能力。不同于传统仅评估单次工具调用的基准,过场生成需数十个相互依赖的工具调用在严格顺序下完成长时程协同——这一能力维度现有基准尚未覆盖。我们在CutsceneBench上评估多种大模型,分析其在此挑战性任务中的表现。

原文摘要 · Abstract (English)

Cutscenes are carefully choreographed cinematic sequences embedded in video games and interactive media, serving as the primary vehicle for narrative delivery, character development, and emotional engagement. Producing cutscenes is inherently complex: it demands seamless coordination across screenwriting, cinematography, character animation, voice acting, and technical direction, often requiring days to weeks of collaborative effort from multidisciplinary teams to produce minutes of polished content. In this work, we present Cutscene Agent, an LLM agent framework for automated end-to-end cutscene generation. The framework makes three contributions: (1)~a Cutscene Toolkit built on the Model Context Protocol (MCP) that establishes \emph{bidirectional} integration between LLM agents and the game engine -- agents not only invoke engine operations but continuously observe real-time scene state, enabling closed-loop generation of editable engine-native cinematic assets; (2)~a multi-agent system where a director agent orchestrates specialist subagents for animation, cinematography, and sound design, augmented by a visual reasoning feedback loop for perception-driven refinement; and (3)~CutsceneBench, a hierarchical evaluation benchmark for cutscene generation. Unlike typical tool-use benchmarks that evaluate short, isolated function calls, cutscene generation requires long-horizon, multi-step orchestration of dozens of interdependent tool invocations with strict ordering constraints -- a capability dimension that existing benchmarks do not cover. We evaluate a range of LLMs on CutsceneBench and analyze their performance across this challenging task.

过场动画多智能体大模型应用游戏开发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。