构建首个基于循环导航的时空一致性评测基准,填补世界模型评估空白
LoopNav: Benchmarking Spatial Consistency in World Models
- 设计循环导航任务,通过真实世界环境数据验证空间一致性
- 采集2000万帧视频,构建包含动作信息的250小时高质量数据集
- 提出场景图一致性评分,可量化空间一致性且不受像素变化干扰
空间一致性是高效世界模型的关键要求,它不仅支持高质量视觉生成,还保障了模拟与规划等下游任务的可靠性。模型需长期保留观测信息,并构建显式或隐式的内部空间表征。然而现有数据集未明确施加空间一致性约束,限制了该能力的系统性评估和数据驱动学习。多数基准侧重视觉连贯性或生成质量,忽视长程空间一致性。为此,我们提出LoopNav,一个以循环导航为核心的全新数据集与评测基准。数据集包含来自Minecraft开放世界中多样化地点的250小时(2000万帧)循环导航视频及对应动作信息。我们进一步引入场景图一致性评分(Scene Graph Consistency Score),在保持对像素级变化不变的同时,量化空间一致性。数据集、评测框架与代码均已开源,旨在推动未来研究。
原文摘要 · Abstract (English)
The ability to simulate the world in a spatially consistent manner is a crucial requirement for effective world models. Such a model enables high-quality visual generation, and also ensures the reliability of world models for downstream tasks such as simulation and planning. It must not only retain long-horizon observational information, but also enables the construction of explicit or implicit internal spatial representations. However, existing datasets do not explicitly enforce spatial consistency constraints, limiting both the ability to systematically evaluate this capability and to learn it through data-driven approaches. Furthermore, most existing benchmarks primarily emphasize visual coherence or generation quality, neglecting the requirement of long-range spatial consistency. To bridge this gap, we propose LoopNav, a dataset and corresponding benchmark centered on loop-based navigation for evaluating spatial consistency. The dataset comprises 250 hours (20 million frames) of loop-based navigation videos with actions, collected from diverse locations in the open-world environment of Minecraft. We further introduce a Scene Graph Consistency Score to quantify spatial consistency while remaining invariant to pixel-level variations. Dataset, benchmark, and code are open-sourced to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。