构建首个跨领域多模态4D世界建模数据集,推动通用场景理解发展。
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- 整合合成与真实数据,覆盖多领域动态交互场景。
- 相较现有数据集规模更大、模态更丰富,支持复杂4D重建与预测任务。
- 适合研究4D建模、视频生成及物理世界理解的团队使用。
4D世界建模旨在同步捕捉空间几何与时间动态,近年来受益于大规模生成模型和多模态学习的发展取得显著进展。然而,通用4D世界模型的发展仍受限于高质量数据的缺乏。现有数据集在动态复杂性、跨域多样性及时空标注方面存在不足,难以支撑4D几何重建、未来预测与相机控制视频生成等关键任务。为此,我们提出OmniWorld——一个大规模、跨领域、多模态的4D世界建模数据集,包含新收集的OmniWorld-Game数据集及多个精选公开数据集。相较于现有合成数据集,OmniWorld-Game具备更丰富的模态覆盖、更大规模和更真实的动态交互。基于此,我们建立了一个具有挑战性的基准,揭示了当前最先进(SOTA)方法在复杂4D环境建模中的局限性。此外,在OmniWorld上微调SOTA模型可显著提升4D重建与视频生成性能,充分验证其作为训练与评估资源的价值。我们期望OmniWorld能加速通用4D世界模型的发展,推动机器对物理世界的整体理解。
原文摘要 · Abstract (English)
The field of 4D world modeling - aiming to jointly capture spatial geometry and temporal dynamics - has witnessed remarkable progress in recent years, driven by advances in large-scale generative models and multimodal learning. However, the development of truly general 4D world models remains fundamentally constrained by the availability of high-quality data. Existing datasets and benchmarks often lack the dynamic complexity, multi-domain diversity, and spatial-temporal annotations required to support key tasks such as 4D geometric reconstruction, future prediction, and camera-control video generation. To address this gap, we introduce OmniWorld, a large-scale, multi-domain, multi-modal dataset specifically designed for 4D world modeling. OmniWorld consists of a newly collected OmniWorld-Game dataset and several curated public datasets spanning diverse domains. Compared with existing synthetic datasets, OmniWorld-Game provides richer modality coverage, larger scale, and more realistic dynamic interactions. Based on this dataset, we establish a challenging benchmark that exposes the limitations of current state-of-the-art (SOTA) approaches in modeling complex 4D environments. Moreover, fine-tuning existing SOTA methods on OmniWorld leads to significant performance gains across 4D reconstruction and video generation tasks, strongly validating OmniWorld as a powerful resource for training and evaluation. We envision OmniWorld as a catalyst for accelerating the development of general-purpose 4D world models, ultimately advancing machines' holistic understanding of the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。