用视觉语言模型评估世界模型生成视频的合理性,提升自动评测精度。
Adapting Vision-Language Models for Evaluating World Models
- 设计多任务评估协议,支持动作与角色识别的多种答题形式。
- 在超5154天的训练中验证方法,在不同条件下达到与专用模型相当的效果。
- 轻量适配且符合人类判断,适合需要语义理解的仿真系统评测。
世界模型——基于历史观测与动作生成环境动态的生成模型——在规划、仿真和具身智能中日益重要。然而,其生成轨迹的评估仍面临挑战,需对动作对齐与语义一致性进行细粒度、时间定位的判断,现有指标难以覆盖。视觉语言模型(VLM)因具备强大的多模态推理能力,有望成为生成内容的自动评估工具,但在细粒度、时间敏感任务中的应用仍有限,需针对性适配。本文提出针对动作识别与角色识别的双重评估任务,涵盖二选一、多选与开放回答三种格式。为支持该任务,我们构建了UNIVERSE(UNIfied Vision-language Evaluator for Rollouts in Simulated Environments),一个在数据与算力受限下适配的VLM评估器。通过超过5,154 GPU天的实验,系统探索了全量、部分及参数高效适配方法,覆盖多种任务格式、上下文长度、采样策略与数据构成。结果表明,统一评估器性能可媲美专用检查点。七种不同环境的人类实验验证其与人工判断高度一致,确立了UNIVERSE作为轻量、可扩展、语义感知的视频世界模型评估新标准。
原文摘要 · Abstract (English)
World models - generative models that simulate environment dynamics conditioned on past observations and actions - are gaining prominence in planning, simulation, and embodied AI. However, evaluating their rollouts remains a fundamental challenge, requiring fine-grained, temporally grounded assessment of action alignment and semantic consistency - capabilities not captured by existing metrics. Vision-Language Models (VLMs) have shown promise as automatic evaluators of generative content due to their strong multimodal reasoning abilities. Yet, their use in fine-grained, temporally sensitive evaluation tasks remains limited and requires targeted adaptation. We introduce an evaluation protocol targeting two recognition tasks - action recognition and character recognition - each assessed across binary, multiple-choice, and open-ended formats. To support this, we present UNIVERSE (UNIfied Vision-language Evaluator for Rollouts in Simulated Environments), a VLM-based evaluator for video world model rollouts adapted under data and compute constraints. In our extensive experiments totaling over 5,154 GPU-days, we explore full, partial, and parameter-efficient adaptation methods across various task formats, context lengths, sampling methods, and data compositions. The resulting unified evaluator achieves parity with task-specific checkpoints. Human studies across seven diverse environments confirm strong alignment with human judgments, establishing UNIVERSE as a lightweight, adaptable, and semantics-aware evaluator for video world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。