评测智能体世界模型的场景、动作与语义质量,填补评估空白
EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
- 构建多维度评测框架,从视觉、运动、语义三方面评估生成效果
- 使用多样化场景与动作模式数据集,揭示现有模型在物理合理性上的不足
- 适合研究具身智能、视频生成与模型评估的学者使用
近年来,创意AI的发展推动了高保真图像和视频的文本条件生成。在此基础上,文本到视频的扩散模型已演变为具备生成物理合理场景能力的具身世界模型(EWMs),有效连接具身智能中的视觉与动作。本文针对现有评估方法仅依赖通用感知指标的不足,提出具身世界模型基准(EWMBench),从视觉场景一致性、运动正确性与语义对齐三个核心维度评估EWMs。该框架基于精心构建的数据集,涵盖多样化的场景与运动模式,并配备全面的多维度评估工具,可系统比较不同模型表现。实验表明,当前视频生成模型在满足具身任务需求方面仍存在明显缺陷。相关数据集与评估工具已在GitHub开源:https://github.com/AgibotTech/EWMBench。
原文摘要 · Abstract (English)
Recent advances in creative AI have enabled the synthesis of high-fidelity images and videos conditioned on language instructions. Building on these developments, text-to-video diffusion models have evolved into embodied world models (EWMs) capable of generating physically plausible scenes from language commands, effectively bridging vision and action in embodied AI applications. This work addresses the critical challenge of evaluating EWMs beyond general perceptual metrics to ensure the generation of physically grounded and action-consistent behaviors. We propose the Embodied World Model Benchmark (EWMBench), a dedicated framework designed to evaluate EWMs based on three key aspects: visual scene consistency, motion correctness, and semantic alignment. Our approach leverages a meticulously curated dataset encompassing diverse scenes and motion patterns, alongside a comprehensive multi-dimensional evaluation toolkit, to assess and compare candidate models. The proposed benchmark not only identifies the limitations of existing video generation models in meeting the unique requirements of embodied tasks but also provides valuable insights to guide future advancements in the field. The dataset and evaluation tools are publicly available at https://github.com/AgibotTech/EWMBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。