用几何信息提升视频生成的长期一致性与物理合理性
DeepVerse: 4D Autoregressive Video Generation as a World Model
- 在生成视频时显式预测并利用几何结构,而非仅依赖视觉
- 相比传统方法,显著减少误差累积和时间不一致,生成更长序列
- 适合需要高保真、长期一致性视频生成的研究与应用
世界模型是实现通用人工智能(AGI)的关键组件,使智能体能通过模拟复杂物理交互来预测未来状态并规划行动。然而,现有交互模型主要预测视觉观测结果,忽视了诸如几何结构和空间连贯性等关键隐藏状态,导致误差快速累积和时间不一致。为解决这一问题,我们提出 DeepVerse——一种新型的4D交互式世界模型,显式将先前时刻的几何预测纳入当前预测,以动作作为条件。实验表明,通过引入显式几何约束,DeepVerse 能捕捉更丰富的时空关系与底层物理动态,显著降低漂移,增强时间一致性,从而可靠生成更长的未来序列,并在预测精度、视觉真实性和场景合理性方面取得显著提升。此外,该方法还提供了有效的几何感知记忆检索方案,有效保持长期空间一致性。我们在多种场景下验证了 DeepVerse 的有效性,证明其具备基于几何感知动力学的高保真、长时程预测能力。
原文摘要 · Abstract (English)
World models serve as essential building blocks toward Artificial General Intelligence (AGI), enabling intelligent agents to predict future states and plan actions by simulating complex physical interactions. However, existing interactive models primarily predict visual observations, thereby neglecting crucial hidden states like geometric structures and spatial coherence. This leads to rapid error accumulation and temporal inconsistency. To address these limitations, we introduce DeepVerse, a novel 4D interactive world model explicitly incorporating geometric predictions from previous timesteps into current predictions conditioned on actions. Experiments demonstrate that by incorporating explicit geometric constraints, DeepVerse captures richer spatio-temporal relationships and underlying physical dynamics. This capability significantly reduces drift and enhances temporal consistency, enabling the model to reliably generate extended future sequences and achieve substantial improvements in prediction accuracy, visual realism, and scene rationality. Furthermore, our method provides an effective solution for geometry-aware memory retrieval, effectively preserving long-term spatial consistency. We validate the effectiveness of DeepVerse across diverse scenarios, establishing its capacity for high-fidelity, long-horizon predictions grounded in geometry-aware dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。