arXiv:2503.16437cs.HCcs.AI2025-03

用文字解谜游戏对比人与大模型的推理能力,发现人类明显更优。

Haunted House: A text-based game for comparing the flexibility of mental models in humans and LLMs

  • 设计九宫格文本迷宫,通过语音提示引导逃脱
  • 人类成功率31.6%,7个大模型仅1次成功(Claude 3 Opus)
  • 大模型常犯随机或不合逻辑的错误,适合评估推理能力

本研究提出名为「Haunted House」的新文本类游戏,用于对比人类与大语言模型(LLMs)在基于模型推理方面的能力。玩家需在3×3网格布局的九间房中避开鬼魂并成功逃脱,每次移动后获得口头线索。研究1中,98名人类参与者取得31.6%的成功率,显著优于测试的七个先进大模型。在七种模型共140次尝试中,仅有一次成功,由Claude 3 Opus完成。初步结果显示GPT o3-mini-high表现较高,但仍未达人类水平。研究2对29名人类参与者移动行为的分析表明,大模型频繁出现随机或非逻辑动作,而人类此类错误较少。研究结果表明,当前大模型在需要主动建模推理的任务中仍存困难,为未来基准测试提供启示。

原文摘要 · Abstract (English)

This study introduces "Haunted House" a novel text-based game designed to compare the performance of humans and large language models (LLMs) in model-based reasoning. Players must escape from a house containing nine rooms in a 3x3 grid layout while avoiding the ghost. They are guided by verbal clues that they get each time they move. In Study 1, the results from 98 human participants revealed a success rate of 31.6%, significantly outperforming seven state-of-the-art LLMs tested. Out of 140 attempts across seven LLMs, only one attempt resulted in a pass by Claude 3 Opus. Preliminary results suggested that GPT o3-mini-high performance might be higher, but not at the human level. Further analysis of 29 human participants' moves in Study 2 indicated that LLMs frequently struggled with random and illogical moves, while humans exhibited such errors less frequently. Our findings suggest that current LLMs encounter difficulties in tasks that demand active model-based reasoning, offering inspiration for future benchmarks.

推理能力人类对比文本游戏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。