arXiv:2503.16068cs.CV2025-03CVPR被引 7

让视频生成更懂物体3D姿态变化,精准跟随旋转轨迹。

PoseTraj: Pose-Aware Trajectory Control in Video Diffusion

  • 用两阶段预训练+3D框监督提升模型对物体姿态的理解。
  • 在10万条合成轨迹数据上训练,实现高精度3D对齐运动生成。
  • 适合需要精准控制物体运动轨迹的视频生成研究者。

轨迹引导的视频生成虽有进展,但在宽范围旋转下难以准确生成具有变化6D姿态的物体运动,根源在于3D理解不足。为此,我们提出PoseTraj,一种基于姿态感知的视频拖拽模型,可从2D轨迹生成3D对齐的运动。方法采用新颖的两阶段姿态感知预训练框架,通过构建大规模合成数据集PoseTraj-10K(含10,000条物体沿旋转轨迹运动的视频),并引入3D边界框作为中间监督信号,增强模型对姿态变化的感知能力。随后,在真实视频上微调轨迹控制模块,并加入相机解耦模块进一步提升运动精度。在多个基准数据集上的实验表明,该方法不仅在旋转轨迹的3D姿态对齐拖拽中表现优异,且在轨迹精度与视频质量上均优于现有基线。

原文摘要 · Abstract (English)

Recent advancements in trajectory-guided video generation have achieved notable progress. However, existing models still face challenges in generating object motions with potentially changing 6D poses under wide-range rotations, due to limited 3D understanding. To address this problem, we introduce PoseTraj, a pose-aware video dragging model for generating 3D-aligned motion from 2D trajectories. Our method adopts a novel two-stage pose-aware pretraining framework, improving 3D understanding across diverse trajectories. Specifically, we propose a large-scale synthetic dataset PoseTraj-10K, containing 10k videos of objects following rotational trajectories, and enhance the model perception of object pose changes by incorporating 3D bounding boxes as intermediate supervision signals. Following this, we fine-tune the trajectory-controlling module on real-world videos, applying an additional camera-disentanglement module to further refine motion accuracy. Experiments on various benchmark datasets demonstrate that our method not only excels in 3D pose-aligned dragging for rotational trajectories but also outperforms existing baselines in trajectory accuracy and video quality.

视频生成扩散模型姿态感知轨迹控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。