用4D几何控制生成动态逼真视频,精准操控相机与多物体运动。
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
- 提出4D几何控制表示,用点云和3D高斯轨迹编码物体运动路径。
- 在35K样本数据集上训练,生成视频视觉质量更高且运动更精准。
- 适合需要精确控制复杂场景动态的视频生成研究者。
视频世界模型旨在模拟动态真实环境,但现有方法难以统一精确控制相机与多物体运动,因视频本质上捕捉的是投影到2D图像平面的动力学。为此,我们提出VerseCrafter,一种基于几何驱动的视频世界模型,可从统一的4D几何世界状态生成动态、逼真的视频。其核心是新型4D几何控制表示,将世界状态编码为静态背景点云和每个物体的3D高斯轨迹。该表示捕捉物体运动路径及随时间变化的概率3D占据,提供灵活且类别无关的替代方案,优于刚性边界框与参数化模型。我们将4D几何控制渲染为4D控制图,输入预训练视频扩散模型,实现高保真、视角一致的视频生成,忠实遵循指定动态。为支持大规模训练,我们构建自动数据引擎,并创建了包含35,000个训练样本的真实世界数据集VerseControl4D,含自动生成的提示和渲染的4D控制图。大量实验表明,VerseCrafter在视觉质量和相机/多物体运动控制精度上均优于现有方法。
原文摘要 · Abstract (English)
Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as videos inherently capture dynamics in the projected 2D image plane. To bridge this gap, we introduce VerseCrafter, a geometry-driven video world model that generates dynamic, realistic videos from a unified 4D geometric world state. Our approach is centered on a novel 4D Geometric Control representation, which encodes the world state as a static background point cloud and per-object 3D Gaussian trajectories. This representation captures each object's motion path and probabilistic 3D occupancy over time, providing a flexible, category-agnostic alternative to rigid bounding boxes and parametric models. We render 4D Geometric Control into 4D control maps for a pretrained video diffusion model, enabling high-fidelity, view-consistent video generation that faithfully follows the specified dynamics. To enable training at scale, we develop an automatic data engine and construct VerseControl4D, a real-world dataset of 35K training samples with automatically derived prompts and rendered 4D control maps. Extensive experiments show that VerseCrafter achieves superior visual quality and more accurate control over camera and multi-object motion than prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。