arXiv:2608.16859cs.CV2026-08被引 1

用智能体自动评估视觉世界模型,让评分有理可查。

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

论文配图:HarnessEval-W: Agentifying the Evaluation of Visual Worlds
图 1 · 摘自论文原文
  • 用多智能体分解问题,逐项诊断生成结果
  • 18个模型330个案例,评分与人类偏好高度一致
  • 公开可扩展的评估框架,适合研究者持续共建

评估不应仅给出一个分数:真正可信的评估在于能解释分数的推理过程。这对世界模型尤为重要,因其需要判断物理、因果和状态演化是否合理。人类能自然识别这些错误,但现有基准无法自动化该能力——指标计算依赖硬编码规则,缺乏可审查的推理链。我们提出HarnessEval-W,将大语言模型领域的‘评测夹具’范式引入世界模型评估。不使用固定标准,而是根据每项评估上下文,将问题分解为可测量子任务,并生成具备特定上下文与诊断工具的专用子智能体分别推理。主智能体验证证据并生成最终结论。这一分层流程使每次评估形成透明的证据树,完整呈现推理链条。我们在330个评估案例中测试了18个代表性世界模型,其判断与人类偏好高度一致,同时提供可验证的细粒度诊断。我们开源完整管道作为动态基准,欢迎社区持续贡献新技能与评估案例以推动世界模型发展。

原文摘要 · Abstract (English)

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.

世界模型智能体评估可解释性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。