构建多轮交互视频世界模型评估基准,覆盖五大核心能力。
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation

- 设计五维评估体系,涵盖画质、设定遵循等五个维度。
- 含289个测试用例、1058轮交互,支持多种视角与操作类型。
- 引入22项自动指标,结果经人工判断验证,适合模型诊断与对比。
交互式世界模型发展迅速,但现有评测基准仅覆盖部分能力,缺乏统一标准。为此,我们提出WBench,一个全面的多轮交互视频世界模型评估基准,涵盖视频质量、场景遵循、交互遵循、一致性及物理合理性五个维度。WBench包含289个测试用例和1058轮交互,每项案例指定世界设定与多轮交互序列,覆盖多样场景、风格、主体及第一/第三人称视角,并支持导航、主体动作、事件编辑、视角切换四类交互。导航支持文本、6-DoF位姿与离散动作控制,兼容不同输入接口。评估采用22项自动子指标,结合专业视觉模型与大模态模型,所有指标均经人类判断验证。对20个前沿模型的评估显示,无模型在所有维度表现优异。我们提供了各模型特征优势、短板及开放挑战的详细诊断分析。代码与数据已开源:https://github.com/meituan-longcat/WBench。
原文摘要 · Abstract (English)
Interactive world models are advancing rapidly, yet existing benchmarks cover only part of the required competencies, leaving no unified standard for systematic evaluation. To fill this gap, we introduce WBench, a comprehensive multi-turn benchmark for interactive world model evaluation along five dimensions, namely video quality, setting adherence, interaction adherence, consistency, and physics compliance. WBench contains 289 test cases and 1,058 interaction turns, where each case specifies a world setting and a multi-turn interaction sequence, covering diverse scenes, styles, subjects, and both first- and third-person perspectives, together with four interaction types, including navigation, subject action, event editing, and perspective switching. For navigation, WBench unifies text, 6-DoF pose, and discrete-action control, enabling evaluation of models with different native input interfaces. Evaluation uses 22 automatic sub-metrics that combine specialist vision models with large multimodal models, and all metrics are validated against human judgments. Across 20 state-of-the-art models, we find that no single model performs strongly across all dimensions. We provide detailed diagnostic insights into the characteristic strengths, weaknesses, and open challenges of each model. Code and data are available at https://github.com/meituan-longcat/WBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。