arXiv:2501.01425cs.CV2025-01ICCV被引 12

实现视频中相机与物体6自由度运动的精确控制

Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation

  • 构建包含完整6D姿态标注的合成数据集SynFMC
  • 可独立或同步控制相机与物体的6D姿态生成高清视频
  • 兼容多种文本到图像模型,适合个性化内容生成

在生成视频中同时控制动态物体和摄像机的运动是一个重要但极具挑战的任务。由于缺乏具备完整6D姿态标注的数据集,现有文本到视频方法无法以3D感知方式协同控制相机与物体的运动,导致生成内容的可控性受限。为解决这一问题并推动该领域研究,我们提出了一个用于自由运动控制的合成数据集SynFMC。该数据集涵盖多样化的物体与环境类别,遵循特定规则覆盖多种运动模式,模拟常见及复杂的现实场景。完整的6D姿态信息使模型能够有效解耦物体与摄像机的运动影响。为进一步实现精准的3D感知运动控制,我们基于SynFMC提出Free-Form Motion Control(FMC)方法。FMC可独立或联合控制物体与摄像机的6D姿态,生成高保真视频,并兼容多种个性化文本到图像(T2I)模型以适应不同内容风格。大量实验表明,所提出的FMC在多个场景下均优于现有方法。

原文摘要 · Abstract (English)

Controlling the movements of dynamic objects and the camera within generated videos is a meaningful yet challenging task. Due to the lack of datasets with comprehensive 6D pose annotations, existing text-to-video methods can not simultaneously control the motions of both camera and objects in 3D-aware manner, resulting in limited controllability over generated contents. To address this issue and facilitate the research in this field, we introduce a Synthetic Dataset for Free-Form Motion Control (SynFMC). The proposed SynFMC dataset includes diverse object and environment categories and covers various motion patterns according to specific rules, simulating common and complex real-world scenarios. The complete 6D pose information facilitates models learning to disentangle the motion effects from objects and the camera in a video.~To provide precise 3D-aware motion control, we further propose a method trained on SynFMC, Free-Form Motion Control (FMC). FMC can control the 6D poses of objects and camera independently or simultaneously, producing high-fidelity videos. Moreover, it is compatible with various personalized text-to-image (T2I) models for different content styles. Extensive experiments demonstrate that the proposed FMC outperforms previous methods across multiple scenarios.

视频生成6D姿态运动控制3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。