arXiv:2409.17331cs.CV2024-09NeurIPS被引 19

用对话控制摄像机运动,让AI像专业摄影师一样拍视频。

ChatCam: Empowering Camera Control through Conversational AI

  • 基于GPT的模型根据文字指令生成摄像机轨迹。
  • 用户对话可精准控制镜头移动,支持复杂拍摄指令。
  • 适合影视制作、虚拟拍摄等需要智能摄像控制的场景。

擅长捕捉世界精髓的电影摄影师通过复杂的摄像机运动创作引人入胜的视觉叙事。随着大型语言模型在感知和交互3D世界方面取得进展,本研究探索其通过人类语言指导控制摄像机的能力。我们提出ChatCam系统,通过与用户的对话导航摄像机运动,模拟专业电影摄影师的工作流程。为此,我们设计CineGPT——一种基于GPT的自回归模型,用于生成文本条件下的摄像机轨迹;同时开发了锚点确定器(Anchor Determinator),以确保轨迹定位精确。ChatCam能够理解用户请求,并利用所提工具生成轨迹,可在辐射场表示上渲染高质量视频。实验包括与最先进方法的对比及用户研究,验证了该方法对复杂摄像指令的解析与执行能力,展现出在真实生产环境中的应用潜力。

原文摘要 · Abstract (English)

Cinematographers adeptly capture the essence of the world, crafting compelling visual narratives through intricate camera movements. Witnessing the strides made by large language models in perceiving and interacting with the 3D world, this study explores their capability to control cameras with human language guidance. We introduce ChatCam, a system that navigates camera movements through conversations with users, mimicking a professional cinematographer's workflow. To achieve this, we propose CineGPT, a GPT-based autoregressive model for text-conditioned camera trajectory generation. We also develop an Anchor Determinator to ensure precise camera trajectory placement. ChatCam understands user requests and employs our proposed tools to generate trajectories, which can be used to render high-quality video footage on radiance field representations. Our experiments, including comparisons to state-of-the-art approaches and user studies, demonstrate our approach's ability to interpret and execute complex instructions for camera operation, showing promising applications in real-world production settings.

视频生成对话控制摄像机轨迹AI导演

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。