构建交互式世界模型评估基准,统一测试感知、记忆与动作生成能力。
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework

- 提出统一动作生成框架,兼容不同交互模态的模型评估。
- 构建33万视频片段数据集,筛选2100个高质量样本覆盖多样场景。
- 设计六类任务共4900个测试样例,全面评估视觉生成与轨迹追踪性能。
实现通用人工智能(AGI)需要具备自适应学习与交互能力的智能体,而交互式世界模型可为感知、推理与行动提供可扩展的环境。然而,当前研究仍缺乏大规模数据集和统一的评估基准来衡量其物理交互能力。为此,我们提出iWorld-Bench,一个用于训练与测试世界模型在距离感知、记忆等交互能力上的综合性基准。我们构建了一个包含33万段视频片段的多样化数据集,并从中精选2100个高质量样本,覆盖多种视角、天气与场景。由于现有世界模型在交互模态上存在差异,我们引入动作生成框架以统一评估标准,并设计六种任务类型,生成4900个测试样本。这些任务共同评估模型在视觉生成、轨迹跟随与记忆能力方面的表现。对14个代表性世界模型的评估揭示了关键局限性,并为未来研究提供了洞见。iWorld-Bench排行榜已公开于iWorld-Bench.com。
原文摘要 · Abstract (English)
Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable environments for perception, reasoning, and action. Yet current research still lacks large-scale datasets and unified benchmarks to evaluate their physical interaction capabilities. To address this, we propose iWorld-Bench, a comprehensive benchmark for training and testing world models on interaction-related abilities such as distance perception and memory. We construct a diverse dataset with 330k video clips and select 2.1k high-quality samples covering varied perspectives, weather, and scenes. As existing world models differ in interaction modalities, we introduce an Action Generation Framework to unify evaluation and design six task types, generating 4.9k test samples. These tasks jointly assess model performance across visual generation, trajectory following, and memory. Evaluating 14 representative world models, we identify key limitations and provide insights for future research. The iWorld-Bench model leaderboard is publicly available at iWorld-Bench.com.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。