arXiv:2504.14128cs.AIcs.CL2025-04被引 11

TALES构建多样文本冒险游戏,评估大模型的复杂推理能力。

TALES: Text Adventure Learning Environment Suite

  • 设计合成与真人撰写的文本冒险游戏挑战模型推理
  • 顶级大模型在人类设计关卡中仅达15%成功率
  • 适合研究模型长程决策与上下文理解的学者

推理是使大型语言模型(LLMs)能够与世界交互的关键能力。随着任务复杂度提升,对顺序决策所需的复杂且多样的推理能力要求越来越高,需基于上下文历史进行结构化推理以决定下一步最佳行动。我们提出TALES,一个包含合成和真人撰写文本冒险游戏的多样化集合,用于挑战和评估多样化的推理能力。我们在一系列LLMs(包括开源与闭源权重模型)上展示结果,并对表现最好的模型进行定性分析。尽管在合成游戏中表现优异,即使是表现最优的LLM代理,在专为人类娱乐设计的游戏上也未能达到15%的成功率。代码与实验可视化详见 https://microsoft.github.io/tale-suite。

原文摘要 · Abstract (English)

Reasoning is an essential skill to enable Large Language Models (LLMs) to interact with the world. As tasks become more complex, they demand increasingly sophisticated and diverse reasoning capabilities for sequential decision-making, requiring structured reasoning over the context history to determine the next best action. We introduce TALES, a diverse collection of synthetic and human-written text-adventure games designed to challenge and evaluate diverse reasoning capabilities. We present results over a range of LLMs, open- and closed-weights, performing a qualitative analysis on the top performing models. Despite an impressive showing on synthetic games, even the top LLM-driven agents fail to achieve 15% on games designed for human enjoyment. Code and visualization of the experiments can be found at https://microsoft.github.io/tale-suite.

文本生成推理能力游戏评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。