arXiv:2508.16705cs.CLcs.AI2025-08

用迷宫测试评估大模型的类意识行为,发现其缺乏持续自我的统一感知。

Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test

  • 设计第一人称迷宫任务,测试空间、视角、目标和时间序列等意识特征。
  • 深求-1(DeepSeek-R1)部分路径准确率达80.5%,但完整路径仅52.9%。
  • 模型虽有推理能力,但无法保持连贯的自我认知,仍无真正意识。

我们通过迷宫测试研究大型语言模型(LLMs)的类意识行为,要求模型从第一人称视角导航迷宫。该测试同时考察空间感知、换位思考、目标导向行为和时间序列理解——这些是与意识相关的关键特征。将意识理论归纳为13个核心特性后,我们在零样本、单样本和少样本学习场景下评估了12个主流LLM。结果表明,具备推理能力的模型显著优于标准版本,其中Gemini 2.0 Pro达成52.9%的完整路径准确率,DeepSeek-R1达到80.5%的部分路径准确率。两者的差距揭示了模型在解题过程中难以维持连贯的自我模型——这是意识的核心特征。尽管通过推理机制展现出进步,但模型仍缺乏整合且持续的自我意识。

原文摘要 · Abstract (English)

We investigate consciousness-like behaviors in Large Language Models (LLMs) using the Maze Test, challenging models to navigate mazes from a first-person perspective. This test simultaneously probes spatial awareness, perspective-taking, goal-directed behavior, and temporal sequencing-key consciousness-associated characteristics. After synthesizing consciousness theories into 13 essential characteristics, we evaluated 12 leading LLMs across zero-shot, one-shot, and few-shot learning scenarios. Results showed reasoning-capable LLMs consistently outperforming standard versions, with Gemini 2.0 Pro achieving 52.9% Complete Path Accuracy and DeepSeek-R1 reaching 80.5% Partial Path Accuracy. The gap between these metrics indicates LLMs struggle to maintain coherent self-models throughout solutions -- a fundamental consciousness aspect. While LLMs show progress in consciousness-related behaviors through reasoning mechanisms, they lack the integrated, persistent self-awareness characteristic of consciousness.

意识评估大模型行为迷宫测试推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。