arXiv:2602.08971cs.CVcs.RO2026-02被引 43

构建统一评测基准,同时评估世界模型的感知质量和实际任务能力

WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models

  • 设计多维度评测体系,涵盖视频质量与智能体任务功能
  • 发现视觉效果好不等于任务表现强,存在显著感知-功能鸿沟
  • 提供公开排行榜,助力推进具实用性的具身世界模型发展

尽管世界模型已成为具身智能的核心,使智能体能通过动作条件预测环境动态,但其评估仍呈碎片化。当前评估主要关注感知保真度(如视频生成质量),忽略了这些模型在下游决策任务中的功能性价值。本文提出WorldArena,一个统一基准,系统评估具身世界模型在感知和功能两个维度的表现。该基准通过三个维度评估:16项指标覆盖六个子维度的视频感知质量;基于人类主观评价的具身任务功能评估,包括将世界模型作为数据引擎、策略评估器和动作规划器;并提出EWMScore,将多维性能整合为单一可解释指标。对14个代表性模型的实验揭示显著的感知-功能差距:高视觉质量并不意味着强任务能力。WorldArena基准及公开排行榜已发布于https://world-arena.ai,为推动真正具功能性的具身世界模型提供评估框架。

原文摘要 · Abstract (English)

While world models have emerged as a cornerstone of embodied intelligence by enabling agents to reason about environmental dynamics through action-conditioned prediction, their evaluation remains fragmented. Current evaluation of embodied world models has largely focused on perceptual fidelity (e.g., video generation quality), overlooking the functional utility of these models in downstream decision-making tasks. In this work, we introduce WorldArena, a unified benchmark designed to systematically evaluate embodied world models across both perceptual and functional dimensions. WorldArena assesses models through three dimensions: video perception quality, measured with 16 metrics across six sub-dimensions; embodied task functionality, which evaluates world models as data engines, policy evaluators, and action planners integrating with subjective human evaluation. Furthermore, we propose EWMScore, a holistic metric integrating multi-dimensional performance into a single interpretable index. Through extensive experiments on 14 representative models, we reveal a significant perception-functionality gap, showing that high visual quality does not necessarily translate into strong embodied task capability. WorldArena benchmark with the public leaderboard is released at https://world-arena.ai, providing a framework for tracking progress toward truly functional world models in embodied AI.

具身智能世界模型评测基准感知-功能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。