提出三维度评测框架,检验世界模型在物理、几何和交互上的长期一致性。
WorldOlympiad: Can Your World Model Survive a Triathlon?

- 分物理、几何、交互三轨评估生成视频的合理性
- 现有模型在长时交互中存在明显物理与结构偏差
- 适合研究生成模型、机器人控制与虚拟环境的开发者
我们提出WorldOlympiad,一个用于诊断基于视频的世界模型的基准测试,涵盖物理真实性、几何一致性与交互保真度三个维度。现有基准多关注视觉质量、语义对齐或短期时间连贯性,难以揭示生成视频是否遵循物理规律、保持一致的3D结构以及在长时程中维持可控交互。WorldOlympiad通过三个子任务实现全面评估:物理赛道利用物体分割与多模态大模型判官,检验机械、热现象和材料属性的可解释规则;几何赛道采用高斯泼溅重建,评估结构一致性、跨视角一致性和相机轨迹对齐;交互赛道则检验生成轨迹是否响应复杂动作指令,并在连续视频片段间保持平滑过渡。该基准覆盖游戏、机器人和真实世界视频三大下游场景,涵盖交互控制、具身操作及开放域运动等挑战。实验表明,当前先进模型在物理推理、3D一致性与长时交互方面仍存在显著差距,凸显了对更系统化评估协议的需求。
原文摘要 · Abstract (English)
We introduce WorldOlympiad, a benchmark for diagnosing video-based world models across physical faithfulness, geometric consistency, and interaction fidelity. While existing benchmarks often focus on visual quality, semantic alignment, or short-term temporal coherence, they provide limited insight into whether generated videos obey physical rules, preserve coherent 3D structure, and sustain controllable interactions over long horizons. To address this gap, WorldOlympiad decomposes world-model evaluation into three complementary dimensions. The physical track uses object segmentation and MLLM-as-judge to assess whether generated videos follow interpretable rules in mechanics, thermal phenomena, and material properties. The geometry track reconstructs generated videos with Gaussian splatting and evaluates structural consistency, cross-view coherence, and camera-trajectory alignment. The interaction track assesses whether generated rollouts follow complex action prompts and maintain smooth, coherent transitions across consecutive video chunks. WorldOlympiad further covers three major downstream scenarios, including gaming, robotics, and general real-world videos, capturing diverse challenges from interactive control and embodied manipulation to open-domain motion and camera dynamics. Together, these tracks and scenarios form a scalable and interpretable evaluation suite that exposes failure modes beyond generic video quality. Experiments on state-of-the-art models reveal substantial gaps in physical reasoning, 3D consistency, and long-horizon interaction, underscoring the need for more structured evaluation protocols for generative world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。