arXiv:2512.18832cs.CL2025-12ACL被引 36

大模型能当文本世界的模拟器吗?实验表明它能提升智能体学习效率。

From Word to World: Can Large Language Models be Implicit Text-based World Models?

  • 把语言模型当作状态预测器,构建三层次评估框架
  • 在五种环境中验证,模型可生成有效轨迹并提升智能体性能
  • 适合研究强化学习与世界模型融合的学者,尤其关注行为覆盖场景

代理强化学习越来越依赖经验驱动的扩展,但真实环境难以适应、覆盖有限且难于扩展。世界模型可通过模拟经验提升学习效率,但大语言模型能否可靠承担此角色仍不明确。本文在文本环境中研究该问题,将语言建模重新诠释为交互下的状态预测。提出三层评估框架:(i) 保真度与一致性,(ii) 可扩展性与鲁棒性,(iii) 代理效用。在五个代表性环境中,我们发现充分训练的世界模型能保持连贯隐状态,随数据和模型规模可预测地扩展,并通过动作验证、合成轨迹生成和强化学习预热显著提升代理性能。然而,这些收益高度依赖行为覆盖范围与环境复杂度,明确了世界建模有效支持学习的边界条件。

原文摘要 · Abstract (English)

Agentic reinforcement learning increasingly relies on experience-driven scaling, yet real-world environments remain non-adaptive, limited in coverage, and difficult to scale. World models offer a potential way to improve learning efficiency through simulated experience, but it remains unclear whether large language models can reliably serve this role and under what conditions they meaningfully benefit agents. We study these questions in text-based environments, which provide a controlled setting to reinterpret language modeling as next-state prediction under interaction. We introduce a three-level framework for evaluating LLM-based world models: (i) fidelity and consistency, (ii) scalability and robustness, and (iii) agent utility. Across five representative environments, we find that sufficiently trained world models maintain coherent latent state, scale predictably with data and model size, and improve agent performance via action verification, synthetic trajectory generation, and warm-starting reinforcement learning. Meanwhile, these gains depend critically on behavioral coverage and environment complexity, delineating clear boundry on when world modeling effectively supports agent learning.

大模型世界模型强化学习文本环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。