统一控制视频生成中的相机与人体运动,提升精度与灵活性。
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
- 用未投影点云实现精准相机控制,无需联合标注数据。
- 在复杂镜头和动作上表现优于现有方法,相机控制更稳定。
- 适合需要精细视觉控制的视频生成研究者使用。
相机与人体运动控制在视频生成中被广泛研究,但现有方法通常分别处理,受限于同时具备高质量标注的稀缺数据。为此,我们提出Uni3C——一个统一的3D增强框架,实现对视频生成中相机与人体运动的精确控制。Uni3C包含两项关键贡献:首先,提出可即插即用的控制模块PCDController,利用单目深度生成的未投影点云,结合冻结的视频生成主干网络,实现准确的相机控制。借助点云的强3D先验与视频基础模型的强大能力,PCDController表现出优异泛化性,无论主干是否微调均表现良好,从而允许不同模块在特定领域(如相机或人体运动)独立训练,降低对联合标注数据的依赖。其次,提出推理阶段的联合对齐3D世界引导机制,无缝融合场景点云与SMPL-X人体模型,分别统一相机与人体运动的控制信号。大量实验表明,PCDController在微调主干上驱动相机运动时具有强鲁棒性;Uni3C在相机可控性与人体运动质量上均显著优于现有方法。此外,我们构建了包含挑战性镜头移动与人体动作的定制验证集,验证了方法的有效性。
原文摘要 · Abstract (English)
Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with high-quality annotations for both aspects. To overcome this, we present Uni3C, a unified 3D-enhanced framework for precise control of both camera and human motion in video generation. Uni3C includes two key contributions. First, we propose a plug-and-play control module trained with a frozen video generative backbone, PCDController, which utilizes unprojected point clouds from monocular depth to achieve accurate camera control. By leveraging the strong 3D priors of point clouds and the powerful capacities of video foundational models, PCDController shows impressive generalization, performing well regardless of whether the inference backbone is frozen or fine-tuned. This flexibility enables different modules of Uni3C to be trained in specific domains, i.e., either camera control or human motion control, reducing the dependency on jointly annotated data. Second, we propose a jointly aligned 3D world guidance for the inference phase that seamlessly integrates both scenic point clouds and SMPL-X characters to unify the control signals for camera and human motion, respectively. Extensive experiments confirm that PCDController enjoys strong robustness in driving camera motion for fine-tuned backbones of video generation. Uni3C substantially outperforms competitors in both camera controllability and human motion quality. Additionally, we collect tailored validation sets featuring challenging camera movements and human actions to validate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。