让文字生成视频像导演拍片一样精准控制3D场景
CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation

- 用户用拖拽方式设定物体和相机在3D中的位置与运动
- 生成视频在空间布局和镜头运动上比现有方法更精准
- 适合影视创作、广告设计等需要精细视觉控制的场景
本文提出CineMaster,一种面向电影级文本到视频生成的3D感知可控框架。目标是让用户获得类似专业导演的控制力:精确摆放场景中物体、灵活操控物体与相机在3D空间中的运动、直观控制画面布局。该框架分两阶段运行:第一阶段通过交互式工作流,用户在3D空间中拖动物体边界框并定义相机轨迹,构建3D感知条件信号;第二阶段将生成的深度图、相机轨迹和物体类别标签作为引导,输入文生视频扩散模型,确保生成符合预期内容的视频。为解决真实世界数据中缺乏3D物体运动与相机姿态标注的问题,我们设计了一套自动化数据标注流程,从大规模视频数据中提取3D边界框与相机轨迹。大量定性与定量实验表明,CineMaster显著优于现有方法,在3D感知文本到视频生成方面表现突出。
原文摘要 · Abstract (English)
In this work, we present CineMaster, a novel framework for 3D-aware and controllable text-to-video generation. Our goal is to empower users with comparable controllability as professional film directors: precise placement of objects within the scene, flexible manipulation of both objects and camera in 3D space, and intuitive layout control over the rendered frames. To achieve this, CineMaster operates in two stages. In the first stage, we design an interactive workflow that allows users to intuitively construct 3D-aware conditional signals by positioning object bounding boxes and defining camera movements within the 3D space. In the second stage, these control signals--comprising rendered depth maps, camera trajectories and object class labels--serve as the guidance for a text-to-video diffusion model, ensuring to generate the user-intended video content. Furthermore, to overcome the scarcity of in-the-wild datasets with 3D object motion and camera pose annotations, we carefully establish an automated data annotation pipeline that extracts 3D bounding boxes and camera trajectories from large-scale video data. Extensive qualitative and quantitative experiments demonstrate that CineMaster significantly outperforms existing methods and implements prominent 3D-aware text-to-video generation. Project page: https://cinemaster-dev.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。