发现预训练智能体与世界模型的规模定律,揭示大模型性能提升规律。
Scaling Laws for Pre-training Agents and World Models
- 基于离线数据的生成学习,建立智能体行为与环境建模的规模关系。
- 发现损失与最优模型规模间存在幂律关系,但系数受分词器、任务和架构影响。
- 为模型与数据规模的最优配置提供依据,适合关注规模化智能体的研究者。
具身智能体的性能随模型参数、数据集规模和计算量增加而提升,已在机器人和视频游戏等领域得到验证。当使用离线数据集上的生成学习目标进行预训练时,可实现对智能体行为(模仿学习)或环境(世界建模)的建模。本文更精确地刻画了规模在这些任务中的作用。超越简单的‘越大越好’直觉,我们发现语言建模中出现的同类幂律也存在于世界建模和模仿学习中(如损失与最优模型规模之间的关系)。然而,这些幂律的系数受分词器、任务及架构显著影响,这对模型与数据的最优规模配置具有重要意义。
原文摘要 · Abstract (English)
The performance of embodied agents has been shown to improve by increasing model parameters, dataset size, and compute. This has been demonstrated in domains from robotics to video games, when generative learning objectives on offline datasets (pre-training) are used to model an agent's behavior (imitation learning) or their environment (world modeling). This paper characterizes the role of scale in these tasks more precisely. Going beyond the simple intuition that `bigger is better', we show that the same types of power laws found in language modeling also arise in world modeling and imitation learning (e.g. between loss and optimal model size). However, the coefficients of these laws are heavily influenced by the tokenizer, task \& architecture -- this has important implications on the optimal sizing of models and data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。