用新基准测试大模型生成符号世界模型的能力
Text2World: Benchmarking Large Language Models for Symbolic World Model Generation
- 基于PDDL构建多领域基准,支持执行评估
- 强化学习训练的模型表现最优,但仍有局限
- 适合研究大模型世界建模与智能体协同的学者
近期,利用大语言模型(LLMs)从文本描述中生成符号世界模型引起关注。尽管已有大量研究探索该方向,但以往工作存在评估随机性强、依赖间接指标、领域覆盖有限等问题。为此,我们提出新基准Text2World,基于规划领域定义语言(PDDL),涵盖数百个多样化领域,并采用多维度、执行驱动的评估指标,提升评测可靠性。我们在Text2World上对当前主流LLMs进行评测,发现经过大规模强化学习训练的推理模型表现更优,但即便最佳模型仍表现出世界建模能力不足。基于此,我们进一步探讨了若干增强策略,包括测试时扩展、智能体训练等。我们希望Text2World能成为推动未来大模型作为世界模型研究的重要资源。项目主页见 https://text-to-world.github.io/。
原文摘要 · Abstract (English)
Recently, there has been growing interest in leveraging large language models (LLMs) to generate symbolic world models from textual descriptions. Although LLMs have been extensively explored in the context of world modeling, prior studies encountered several challenges, including evaluation randomness, dependence on indirect metrics, and a limited domain scope. To address these limitations, we introduce a novel benchmark, Text2World, based on planning domain definition language (PDDL), featuring hundreds of diverse domains and employing multi-criteria, execution-based metrics for a more robust evaluation. We benchmark current LLMs using Text2World and find that reasoning models trained with large-scale reinforcement learning outperform others. However, even the best-performing model still demonstrates limited capabilities in world modeling. Building on these insights, we examine several promising strategies to enhance the world modeling capabilities of LLMs, including test-time scaling, agent training, and more. We hope that Text2World can serve as a crucial resource, laying the groundwork for future research in leveraging LLMs as world models. The project page is available at https://text-to-world.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。