arXiv:2608.02603cs.CV2026-08被引 1

评测视频生成模型的内在反应能力,揭示视觉质量不等于世界逻辑自洽。

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

论文配图:WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
图 1 · 摘自论文原文
  • 构建四层分级评测体系,涵盖视觉、控制、空间一致性与世界反应性。
  • 20个模型测试显示:无一同时具备强反应性与广泛任务覆盖能力。
  • 语言驱动模型交互更好,但复杂指令执行常出错,适合对齐用户意图场景。

可控视频生成模型正日益作为世界模型被开发。评估这类模型需超越生成视频的外观表现,考察其内在反应性——即根据场景状态推断世界应如何响应,并生成未在输入中明确描述的合理后果。然而现有基准主要关注视觉质量或显式指令执行情况,忽视了内在反应性的评估。为此,本文提出WorldExam,一个分层诊断性基准,包含四个层级:视觉质量、控制遵循性、空间一致性与世界反应性。该基准涵盖8项专门任务,共1,474个案例,支持对相机、动作和语言驱动模型范式的统一评估。其中,世界反应性层级评估输入场景条件下的情境化反应与目标导向行为。对20个代表性模型的评估发现明显能力分化:相机驱动模型擅长相机控制,但无法实现动态交互;动作驱动模型能更精确地控制主体,但常导致世界无响应;语言驱动模型在交互方面表现更优,但在执行复杂控制时可靠性下降。没有模型能同时在广泛任务上保持强反应性,说明高视觉质量与指令符合度并不保证内在反应性。

原文摘要 · Abstract (English)

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.

世界模型视频生成评估基准反应性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。