arXiv:2605.08578cs.LGcs.AI2026-05

研究大模型规模对游戏世界模型数据效率的影响,发现联合训练可稳定提升性能。

Probing the Impact of Scale on Data-Efficient, Generalist Transformer World Models for Atari

论文配图:Probing the Impact of Scale on Data-Efficient, Generalist Transformer World Models for Atari
图 1 · 摘自论文原文
  • 用极简Transformer世界模型测试不同规模下的表现
  • 26个游戏联合训练后,所有环境均实现性能单调提升
  • 模型精度提升直接转化为下游控制策略的显著改善

开发具备类人数据效率的通用系统是核心挑战。尽管世界模型(WMs)前景可观,但现有研究常将架构机制与模型规模影响混为一谈。本文采用极简Transformer世界模型,在Atari 100k基准上分析规模效应,使用固定离线数据集(来自预设专家策略)。结果表明,不同环境呈现截然不同的规模扩展规律,即使数据预算和模型容量相同。某些任务可越过插值阈值,实现过参数化下持续优化;另一些则困于经典范式,模型越大性能越差。在统一设置下(单个Transformer训练26个Atari环境),联合训练稳定了扩展动态,确保所有环境均获得单调增益。最终,更高精度的模型直接提升下游控制效果:仅在仿真环境中学习的策略,达到中位专家-随机-归一化得分0.770。研究提示,未来进展不仅依赖架构创新,更需精准的规模策略。

原文摘要 · Abstract (English)

Developing generalist systems that retain human-like data efficiency is a central challenge. While world models (WMs) offer a promising path, existing research often conflates architectural mechanisms with the independent impact of model \emph{scale}. In this work, we use a minimalist transformer world model to analyze scaling behaviors on the Atari 100k benchmark, using fixed offline datasets derived from a presupposed expert policy. Our results reveal that environments fundamentally fall into distinct scaling regimes, even when constrained by identical offline data budgets and model capacities. For individual tasks, some environments naturally allow models to pass the interpolation threshold, yielding monotonic improvements in the overparameterized regime, while others remain trapped in the classical regime, where larger world models degrade fidelity. In the unified setting, i.e., a single transformer trained on a suite of 26 Atari environments, we uncover that joint training stabilizes scaling dynamics, ensuring monotonic gains across all environments, regardless of their distinct inherent scaling regimes. Finally, we demonstrate that improved fidelity translates directly to downstream control, with policies learned entirely within the simulated dynamics achieving a median expert-random-normalized score of 0.770. Our findings suggest that future progress lies as much in precise scaling strategies as in architectural innovation.

世界模型规模效应强化学习Atari

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。