让模型看懂4D世界,生成未见时空内容
SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
- 从单视角2D帧重建连续4D动态世界
- 在低秩表示与物理约束下学习4D动态
- 支持未见时空生成与视频编辑,适合视觉生成研究者
图像和视频是4D世界(3D空间+时间)的离散二维投影。现有视觉理解、预测与生成方法大多直接作用于2D观测,导致性能受限。我们提出SeeU,一种新方法,通过学习连续4D动态来生成未见视觉内容。其核心是2D→4D→2D的学习框架:首先从稀疏单视角2D帧重建4D世界(2D→4D);然后在低秩表示与物理约束下学习连续4D动态(离散4D→连续4D);最后向前滚动世界时间,采样时间点与视角重新投影至2D,基于时空上下文生成未见区域(4D→2D)。通过建模4D动态,SeeU实现连续且物理一致的新视觉生成,在未见时间生成、未见空间生成与视频编辑等任务中表现优异。所有数据与代码将公开于https://yuyuanspace.com/SeeU/
原文摘要 · Abstract (English)
Images and videos are discrete 2D projections of the 4D world (3D space + time). Most visual understanding, prediction, and generation operate directly on 2D observations, leading to suboptimal performance. We propose SeeU, a novel approach that learns the continuous 4D dynamics and generate the unseen visual contents. The principle behind SeeU is a new 2D$\to$4D$\to$2D learning framework. SeeU first reconstructs the 4D world from sparse and monocular 2D frames (2D$\to$4D). It then learns the continuous 4D dynamics on a low-rank representation and physical constraints (discrete 4D$\to$continuous 4D). Finally, SeeU rolls the world forward in time, re-projects it back to 2D at sampled times and viewpoints, and generates unseen regions based on spatial-temporal context awareness (4D$\to$2D). By modeling dynamics in 4D, SeeU achieves continuous and physically-consistent novel visual generation, demonstrating strong potentials in multiple tasks including unseen temporal generation, unseen spatial generation, and video editing. All data and code will be public at https://yuyuanspace.com/SeeU/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。