arXiv:2411.18644cs.CV2024-11

用语言+人工干预生成逼真视频,解决动态不一致问题

Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop

  • 用大模型把文字转为3D场景指令,自动构建场景
  • 用户可实时通过界面调整,生成无物理错误的视频
  • 适合需要精细控制场景的影视/游戏制作人员

视频生成虽已取得显著质量提升,但仍存在时间不一致和违反物理规律等伪影问题。借助3D场景可从根本上解决这些问题,提供对场景元素的精确控制。为便于生成多样化的逼真3D场景,我们提出Scene Co-pilot框架,结合大语言模型(LLMs)与程序化3D场景生成器。该框架包含Scene Codex、BlenderGPT和人工介入环节。Scene Codex负责将用户文本输入转化为3D场景生成器可理解的指令;BlenderGPT为用户提供直观直接的方式,精准控制生成的3D场景及最终视频输出,并支持通过Blender UI即时获取视觉反馈。此外,我们还构建了一个以代码形式呈现的程序化物体数据集,进一步增强系统能力。各组件协同工作,支持用户高效生成所需3D场景。大量实验表明,本框架在自定义3D场景和视频生成方面具备强大能力。

原文摘要 · Abstract (English)

Video generation has achieved impressive quality, but it still suffers from artifacts such as temporal inconsistency and violation of physical laws. Leveraging 3D scenes can fundamentally resolve these issues by providing precise control over scene entities. To facilitate the easy generation of diverse photorealistic scenes, we propose Scene Copilot, a framework combining large language models (LLMs) with a procedural 3D scene generator. Specifically, Scene Copilot consists of Scene Codex, BlenderGPT, and Human in the loop. Scene Codex is designed to translate textual user input into commands understandable by the 3D scene generator. BlenderGPT provides users with an intuitive and direct way to precisely control the generated 3D scene and the final output video. Furthermore, users can utilize Blender UI to receive instant visual feedback. Additionally, we have curated a procedural dataset of objects in code format to further enhance our system's capabilities. Each component works seamlessly together to support users in generating desired 3D scenes. Extensive experiments demonstrate the capability of our framework in customizing 3D scenes and video generation.

视频生成3D场景人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。