arXiv:2505.13934cs.LGcs.AI2025-05NeurIPS被引 49

用强化学习直接优化世界模型的预测准确性和感知质量。

RLVR-World: Training World Models with Reinforcement Learning

  • 用可验证奖励的强化学习训练世界模型,对齐任务目标。
  • 在文本游戏、网页导航和机器人操作中显著提升性能。
  • 适合希望提升生成模型实用性的研究者和工程师。

世界模型通过预测动作引发的状态变化而日益发展,但标准训练目标如最大似然估计(MLE)常与任务特定目标(如预测准确率或感知质量)不一致。本文提出RLVR-World框架,利用可验证奖励的强化学习(RLVR),直接优化世界模型以实现这些指标。尽管世界建模被形式化为分词序列的自回归预测,但RLVR-World将解码后预测结果的评估指标作为可验证奖励进行反馈。我们在跨领域的语言和视频世界模型上均取得显著性能提升,涵盖文本游戏、网页导航和机器人操作任务。研究表明,除近期大模型推理能力进步外,RLVR为增强生成模型实用性提供了一种有前景的后训练范式。代码、数据集、模型及视频样例可在项目网站获取:https://thuml.github.io/RLVR-World。

原文摘要 · Abstract (English)

World models predict state transitions in response to actions and are increasingly developed across diverse modalities. However, standard training objectives such as maximum likelihood estimation (MLE) often misalign with task-specific goals of world models, i.e., transition prediction metrics like accuracy or perceptual quality. In this paper, we present RLVR-World, a unified framework that leverages reinforcement learning with verifiable rewards (RLVR) to directly optimize world models for such metrics. Despite formulating world modeling as autoregressive prediction of tokenized sequences, RLVR-World evaluates metrics of decoded predictions as verifiable rewards. We demonstrate substantial performance gains on both language- and video-based world models across domains, including text games, web navigation, and robot manipulation. Our work indicates that, beyond recent advances in reasoning language models, RLVR offers a promising post-training paradigm for enhancing the utility of generative models more broadly. Code, datasets, models, and video samples are available at the project website: https://thuml.github.io/RLVR-World.

世界模型强化学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。