直接生成4D动态3D物体,实现高保真与时间连贯性
SS4D: Native 4D Generative Model via Structured Spacetime Latents
- 用结构化时空潜变量直接建模4D数据,避免逐帧优化
- 在10秒长视频上实现98%的时间一致性与高保真重建
- 适合需要高效生成动态3D内容的研究者与开发者
我们提出SS4D,一种原生4D生成模型,可直接从单目视频合成动态3D物体。不同于以往通过优化3D或视频生成模型构建4D表示的方法,我们直接在4D数据上训练生成器,实现了高保真、时间连贯性和结构一致性。核心在于一组压缩的结构化时空潜变量:(1)为应对4D训练数据稀缺,基于预训练单图像到3D模型,保持强空间一致性;(2)引入专用时序层,跨帧推理以保证时间连贯性;(3)通过因子化4D卷积和时序下采样块,在时间轴上压缩潜变量序列,支持长视频高效训练与推理。此外,采用精心设计的训练策略,增强对遮挡的鲁棒性。
原文摘要 · Abstract (English)
We present SS4D, a native 4D generative model that synthesizes dynamic 3D objects directly from monocular video. Unlike prior approaches that construct 4D representations by optimizing over 3D or video generative models, we train a generator directly on 4D data, achieving high fidelity, temporal coherence, and structural consistency. At the core of our method is a compressed set of structured spacetime latents. Specifically, (1) To address the scarcity of 4D training data, we build on a pre-trained single-image-to-3D model, preserving strong spatial consistency. (2) Temporal consistency is enforced by introducing dedicated temporal layers that reason across frames. (3) To support efficient training and inference over long video sequences, we compress the latent sequence along the temporal axis using factorized 4D convolutions and temporal downsampling blocks. In addition, we employ a carefully designed training strategy to enhance robustness against occlusion
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。