首个统一评估世界生成的基准,涵盖3000个多样化场景。
WorldScore: A Unified Evaluation Benchmark for World Generation

- 将世界生成拆解为带相机轨迹的连续场景生成任务。
- 覆盖19个模型,揭示不同类别模型在可控性、质量与动态上的短板。
- 适合研究3D/4D场景生成与视频生成的学者与开发者使用。
我们提出WorldScore基准,首个统一的世界生成评估框架。将世界生成分解为一系列基于相机轨迹布局的下一场景生成任务,实现对3D/4D场景生成到视频生成等多种方法的统一评估。该基准包含3000个精心筛选的测试样本,覆盖静态与动态、室内外、写实与风格化等多样世界。评估指标从可控性、质量与动态性三个维度衡量生成世界的表现。通过对19个代表性模型(含开源与闭源)的广泛评测,揭示了各类模型的关键挑战与洞察。数据集、评估代码及排行榜可访问 https://haoyi-duan.github.io/WorldScore/
原文摘要 · Abstract (English)
We introduce the WorldScore benchmark, the first unified benchmark for world generation. We decompose world generation into a sequence of next-scene generation tasks with explicit camera trajectory-based layout specifications, enabling unified evaluation of diverse approaches from 3D and 4D scene generation to video generation models. The WorldScore benchmark encompasses a curated dataset of 3,000 test examples that span diverse worlds: static and dynamic, indoor and outdoor, photorealistic and stylized. The WorldScore metrics evaluate generated worlds through three key aspects: controllability, quality, and dynamics. Through extensive evaluation of 19 representative models, including both open-source and closed-source ones, we reveal key insights and challenges for each category of models. Our dataset, evaluation code, and leaderboard can be found at https://haoyi-duan.github.io/WorldScore/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。