arXiv:2605.07390cs.CV2026-05

让生成模型理解4维时空规律,实现更真实的动态场景生成。

ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation

论文配图:ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation
图 1 · 摘自论文原文
  • 构建时空认知图,融合全局结构与局部动态信息。
  • 在3个3D/4D生成任务上超越现有方法,提升时空一致性。
  • 适合研究视频生成、物理模拟的开发者和研究人员。

生成模型在生成2D视频方面取得成功,但在真实物理世界中仍面临4维时空尺度的挑战。现有4D生成模型通常直接嵌入宏观尺度约束以增强整体时空一致性,但仅保证外观连贯性,无法揭示物理世界的局部动态。本文提出基于4维时空认知的世界模型框架ST-Gen4D,通过四个关键设计:1)多模态特征表示;2)构建全局外观图与局部动态图,并通过语义桥接融合成4维认知图;3)利用世界模型推演未来状态;4)以推导出的认知为条件,引导潜空间扩散生成4D高斯。该框架深度融合4维内在认知与生成先验,确保生成结果的结构合理性与拓扑一致性。同时构建了ST-4D数据集,整合公开数据与自建子集。大量实验表明,ST-Gen4D在3个3D与4D生成任务中表现优异。

原文摘要 · Abstract (English)

Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically, existing 4D generative models directly embed macro scale constraints to enhance overall spatiotemporal consistency. However, these methods only ensure global appearance coherence and fail to reveal the local dynamics of the physical world. Our insight is that global appearance structure and local dynamic topology empower 4D spatiotemporal cognition, thereby enabling 4D generation with spatiotemporal regularities. In this work, we propose ST-Gen4D, a 4D generation framework with 4D spatiotemporal cognition-based world model. Our model is guided by four key designs: 1) Spatiotemporal representation. We encode various modalities into multiple representations as a feature basis. 2) Spatiotemporal cognition. We sculpture these representations into global appearance graph and local dynamic graph, and fuse them via semantic-bridged spatiotemporal fusion to obtain a 4D cognition graph. 3) Spatiotemporal reasoning. We utilize a world model to derive future state based on the 4D cognition. 4) Spatiotemporal generation. We leverage the derived cognition as condition to guide latent diffusion for 4D Gaussian generation. By deeply integrating 4D intrinsic cognition with generative priors, our model guarantees the structural rationality and topological consistency of 4D generation. Moreover, we propose ST-4D datasets by aggregating public 4D datasets and self-built subset. Extensive experiments demonstrate the superiority of our ST-Gen4D across 3D and 4D generation tasks.

4D生成时空认知世界模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。