arXiv:2512.03400cs.LGcs.AI2025-12被引 3

用魔方训练模型,显式建模世界能提升性能。

Better World Models Can Lead to Better Post-Training Performance

  • 用状态预测监督增强模型世界建模能力
  • 世界模型质量越高,下游任务表现越好
  • 小部分数据用于预训练效果最佳

我们研究了显式世界建模目标如何影响Transformer的内部表征与下游能力,以魔方为训练场景。探究三个问题:(1) 显式预训练世界模型如何改变模型潜在表征;(2) 世界模型质量对后期训练性能的影响;(3) 在有限数据预算下,如何分配预训练与任务微调的比例。对比标准动作预测微调,以及两种添加显式状态预测监督的策略:先状态预训练再微调,以及动作与状态联合训练。使用组相对策略优化(GRPO)评估任务准确率。发现显式世界建模可提升表征质量,表现为更高的探针准确率和可操控性,且高质量表征在GRPO中带来更大增益,尤其在更复杂的魔方状态上。当微调数据量固定时,探针准确率强烈预测GRPO提升幅度。在总数据量固定的情况下,仅将少量数据用于预训练可最大化准确率。

原文摘要 · Abstract (English)

We study how explicit world-modeling objectives affect the internal representations and downstream capability of Transformers, using Rubik's Cubes as our training domain. We ask: (1) how does explicitly pretraining a world model affect a model's latent representations, (2) how does world-model quality affect post-training performance, and (3) how should a finite data budget be split between pretraining and task-specific fine-tuning? We compare standard action-prediction fine-tuning with two strategies that add explicit state-prediction supervision: state pretraining followed by fine-tuning, and joint action and state training. We measure task accuracy after Group Relative Policy Optimization (GRPO). We further find that explicit world-modeling yields better representations in terms of higher probing accuracy and steerability of the model, and that better representations yield larger gains from GRPO, especially on harder cube states. Finally, when the fine-tuning budget is held fixed, probe accuracy strongly predicts the GRPO improvement. Under a fixed total data budget, accuracy is maximized by allocating only a small fraction to pretraining.

世界模型魔方表征学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。