测试视频世界模型能否在不观察时仍持续演化
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
- 设计可控遮挡、关灯、视线转移等指令测试模型
- 发现多数模型无法脱离观察独立演化
- 揭示当前模型在数据与结构上的潜在偏差
现实中的变化,如水流或冰融化,无论是否被观察都会发生。视频世界模型通过二维帧观测生成“世界”。这些生成的“世界”能否在未被观察时仍持续演化?为探究此问题,我们设计了基准测试 STEVO-Bench,通过插入遮挡物、关闭灯光或指定相机“移开视线”轨迹,对自然演化过程施加观察控制。在多种自然演化场景下,对比有无相机控制的视频模型表现,暴露其难以将状态演化与观察解耦的问题。STEVO-Bench 提出一种自动检测并分离模型在自然状态演化关键方面失败模式的评估协议。对 STEVO-Bench 结果的分析揭示了当前视频世界模型在数据与架构上的潜在偏见。
原文摘要 · Abstract (English)
Evolutions in the world, such as water pouring or ice melting, happen regardless of being observed. Video world models generate "worlds" via 2D frame observations. Can these generated "worlds" evolve regardless of observation? To probe this question, we design a benchmark to evaluate whether video world models can decouple state evolution from observation. Our benchmark, STEVO-Bench, applies observation control to evolving processes via instructions of occluder insertion, turning off the light, or specifying camera "lookaway" trajectories. By evaluating video models with and without camera control for a diverse set of naturally-occurring evolutions, we expose their limitations in decoupling state evolution from observation. STEVO-Bench proposes an evaluation protocol to automatically detect and disentangle failure modes of video world models across key aspects of natural state evolution. Analysis of STEVO-Bench results provide new insight into potential data and architecture bias of present-day video world models. Project website: https://glab-caltech.github.io/STEVOBench/. Blog: https://ziqi-ma.github.io/blog/2026/outofsight/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。