用3D点轨迹增强视频生成,让物体运动更符合物理规律。
Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach
- 通过2D视频添加3D点轨迹,构建3D感知的视频数据集。
- 引入形状与运动正则化,消除非物理形变等伪影。
- 适合需要精准物体交互建模的任务型视频生成场景。
我们提出一种融合三维几何与动态感知的新型视频生成框架。通过在2D视频中引入3D点轨迹并将其对齐至像素空间,构建了名为PointVid的3D感知视频数据集,并用于微调潜在扩散模型,使其能够以3D笛卡尔坐标追踪2D物体。在此基础上,我们对视频中物体的形状与运动施加正则化约束,有效消除非物理形变等不良伪影。实验表明,该方法显著提升了生成视频的视觉质量,缓解了当前视频模型普遍存在的物体变形问题。所提方法可无缝集成至现有视频扩散模型中,尤其适用于包含丰富接触交互的任务导向型视频生成任务,其中3D信息对于理解物体形状与运动至关重要。
原文摘要 · Abstract (English)
We present a novel video generation framework that integrates 3-dimensional geometry and dynamic awareness. To achieve this, we augment 2D videos with 3D point trajectories and align them in pixel space. The resulting 3D-aware video dataset, PointVid, is then used to fine-tune a latent diffusion model, enabling it to track 2D objects with 3D Cartesian coordinates. Building on this, we regularize the shape and motion of objects in the video to eliminate undesired artifacts, e.g., non-physical deformation. Consequently, we enhance the quality of generated RGB videos and alleviate common issues like object morphing, which are prevalent in current video models due to a lack of shape awareness. With our 3D augmentation and regularization, our model is capable of handling contact-rich scenarios such as task-oriented videos, where 3D information is essential for perceiving shape and motion of interacting solids. Our method can be seamlessly integrated into existing video diffusion models to improve their visual plausibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。