arXiv:2511.23465cs.LG2025-11被引 4

构建可控环境评测基准,检验世界模型对动态规律的理解能力。

SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments

  • 设计小型封闭环境基准,无手工奖励信号,精确控制动态变化。
  • 六类场景下测试四类模型,发现长期预测性能随步数显著下降。
  • 适合研究表示学习与动态建模的学者,揭示当前模型局限性。

当前世界模型缺乏统一且受控的评估环境,难以判断其是否真正捕捉到环境动态的底层规律。本文提出SmallWorld基准,一个用于在孤立、精确控制的动态环境下评估世界模型能力的测试平台,不依赖人工设计的奖励信号。基于该基准,我们在完全可观测状态空间中对递归状态空间模型、Transformer、扩散模型和神经微分方程等代表性架构进行了全面实验,覆盖六个不同领域。实验结果揭示了这些模型对环境结构的建模能力及其在长序列滚动中的预测退化现象,凸显了现有建模范式的优缺点,并为表示学习与动态建模的未来发展提供了重要洞见。

原文摘要 · Abstract (English)

Current world models lack a unified and controlled setting for systematic evaluation, making it difficult to assess whether they truly capture the underlying rules that govern environment dynamics. In this work, we address this open challenge by introducing the SmallWorld Benchmark, a testbed designed to assess world model capability under isolated and precisely controlled dynamics without relying on handcrafted reward signals. Using this benchmark, we conduct comprehensive experiments in the fully observable state space on representative architectures including Recurrent State Space Model, Transformer, Diffusion model, and Neural ODE, examining their behavior across six distinct domains. The experimental results reveal how effectively these models capture environment structure and how their predictions deteriorate over extended rollouts, highlighting both the strengths and limitations of current modeling paradigms and offering insights into future improvement directions in representation learning and dynamics modeling.

世界模型动态建模基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。