arXiv:2412.13549cs.CLcs.AI2024-12ACL被引 3

测试大模型在陌生环境中的创造力,发现其解题能力仅15%。

EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents

  • 设计房间逃脱游戏环境,考察模型的创意推理与非常规工具使用能力。
  • 现有模型无提示下平均仅完成15%进度,暴露创造力短板。
  • 提出新框架,让模型能自主规划超千步动作链,少用提示更高效解谜。

语言模型代理在长时规划与推理方面表现优异,但现有基准多聚焦目标明确的任务,忽视了在陌生环境中创造性适应的能力。为此,我们提出EscapeBench,一个由房间逃脱游戏环境组成的基准套件,旨在挑战代理在隐含目标下的创造性推理、非常规工具使用及迭代求解能力。实验表明,尽管当前语言模型采用工作记忆和思维链推理,但在无提示情况下平均仅实现15%的进度,凸显其在创造性方面的局限性。为弥补这一差距,我们提出EscapeAgent框架,通过前瞻(创新性工具使用)与反思(识别未解决任务)增强创造性推理。实验显示,EscapeAgent可执行超过1000步的动作链并保持逻辑连贯性,以最多40%更少的步骤与提示完成游戏,在不同难度下均表现稳健,并以更高成功率实现更高效、更具创新性的解谜策略。

原文摘要 · Abstract (English)

Language model agents excel in long-session planning and reasoning, but existing benchmarks primarily focus on goal-oriented tasks with explicit objectives, neglecting creative adaptation in unfamiliar environments. To address this, we introduce EscapeBench, a benchmark suite of room escape game environments designed to challenge agents with creative reasoning, unconventional tool use, and iterative problem-solving to uncover implicit goals. Our results show that current LM models, despite employing working memory and Chain-of-Thought reasoning, achieve only 15% average progress without hints, highlighting their limitations in creativity. To bridge this gap, we propose EscapeAgent, a framework designed to enhance creative reasoning through Foresight (innovative tool use) and Reflection (identifying unsolved tasks). Experiments show that EscapeAgent can execute action chains over 1,000 steps while maintaining logical coherence. It navigates and completes games with up to 40% fewer steps and hints, performs robustly across difficulty levels, and achieves higher action success rates with more efficient and innovative puzzle-solving strategies.

语言模型创造力智能体推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。