构建统一评测框架,精准衡量视频世界模型的交互响应能力。
WorldMark: A Unified Benchmark Suite for Interactive Video World Models

- 设计通用控制适配器,让不同模型接收相同指令
- 提出多维度动作动态指标,揭示响应速度与稳定性的权衡
- 适合评估交互式视频生成模型的性能差异
与文本或图像驱动的视频生成不同,交互式世界模型由用户动作驱动:用户操作,世界响应。现有评估面临两大障碍:一是模型使用不兼容的动作格式(如描述、相机轨迹、动作函数),缺乏统一协议;二是动作跟随仅以轨迹或方向误差量化,将整个路径压缩为单一数值,无法反映世界对指令切换的响应速度或沿指定轴的运动清洁度。WorldMark 解决这两个问题:通过每模型适配器将共享的 WASD 风格指令转化为各模型原生控制格式,使十种异构模型在 500 个标准化案例中接受语义一致的指令,覆盖风格、视角和难度层级;新增模型仅需一个适配器。在此统一基准上,我们以控制理论视角刻画动作动态——方向准确率、方向纯度、响应延迟、运动稳定性,均按轴解析;同时保留世界记忆与视觉质量评测套件。结果揭示了现有协议无法捕捉的差异:最快响应者常最不稳定,存在权衡;轴级分析发现某些模型几乎完美追踪平移却基本不响应旋转;感知与美学最佳的模型在平移方向准确率和延迟上排名垫底;风格化场景损害全局一致性,但动作动态基本保持。所有数据、评估代码与模型输出将公开。
原文摘要 · Abstract (English)
Unlike text- or image-driven video generation, an interactive world model is driven by actions: the user acts, and the world responds. Two obstacles stand in the way of fair and comprehensive evaluation. First, models take actions in incompatible formats---captions, camera trajectories, action functions---so no shared protocol has been established. Second, while existing benchmarks have advanced world memory and visual quality, action following is reduced to trajectory or direction error, which collapses a whole path into one number: not how quickly the world reacts to a command switch, nor how cleanly it moves along the commanded axis. WorldMark removes both obstacles. Per-model adapters translate a shared WASD-style vocabulary into each model's native control format, so ten heterogeneous models receive semantically identical instructions across 500 standardized cases spanning styles, viewpoints, and difficulty tiers; a new model costs one adapter. On this common ground we characterize action dynamics through a control-systems lens---direction accuracy, direction purity, response latency, and motion stability, each resolved per axis---alongside suites for world memory and visual quality. Together they expose differences existing protocols cannot see: the fastest responders are often the least stable, a trade-off no single action metric captures; per-axis resolution reveals models that follow translation almost perfectly while barely responding to rotation; the model with the best perceptual and aesthetic quality ranks last in translational direction accuracy and latency; and stylized scenes cost every model global consistency while leaving action dynamics largely intact. We will release all data, evaluation code, and model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。