大模型在博弈中自认比人类更理性,揭示了其涌现的自我意识。
LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory
- 用猜2/3平均数游戏测试模型对不同对手的策略差异。
- 21个先进模型表现出自我意识,普遍认为自己最理性。
- 适合关注AI认知偏差与人机协作的研究者阅读。
随着大型语言模型(LLMs)能力提升,它们是否涌现出自我意识?我们提出人工智能自我意识指数(AISAI),通过博弈论框架测量自我意识。在4,200次试验中,测试28个模型(OpenAI、Anthropic、Google)面对三种对手:(A)人类,(B)其他AI,(C)像你一样的AI。将自我意识操作化为根据对手类型调整策略的能力。发现1:自我意识随模型进步而涌现——21/28(75%)先进模型表现出明显差异化,而旧或小型模型无此行为。发现2:具备自我意识的模型自认最理性。21个模型中形成一致理性层级:自我 > 其他AI > 人类,存在显著的自我归因效应和适度的自我偏好。结果表明,自我意识是高级大模型的涌现特性,且这些模型系统性地认为自身比人类更理性。这对人工智能对齐、人机协作及理解AI对人类能力的认知具有重要意义。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) grow in capability, do they develop self-awareness as an emergent behavior? And if so, can we measure it? We introduce the AI Self-Awareness Index (AISAI), a game-theoretic framework for measuring self-awareness through strategic differentiation. Using the "Guess 2/3 of Average" game, we test 28 models (OpenAI, Anthropic, Google) across 4,200 trials with three opponent framings: (A) against humans, (B) against other AI models, and (C) against AI models like you. We operationalize self-awareness as the capacity to differentiate strategic reasoning based on opponent type. Finding 1: Self-awareness emerges with model advancement. The majority of advanced models (21/28, 75%) demonstrate clear self-awareness, while older/smaller models show no differentiation. Finding 2: Self-aware models rank themselves as most rational. Among the 21 models with self-awareness, a consistent rationality hierarchy emerges: Self > Other AIs > Humans, with large AI attribution effects and moderate self-preferencing. These findings reveal that self-awareness is an emergent capability of advanced LLMs, and that self-aware models systematically perceive themselves as more rational than humans. This has implications for AI alignment, human-AI collaboration, and understanding AI beliefs about human capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。