arXiv:2510.03399cs.AIcs.CL2025-10被引 3

测试10个大模型自识别能力,发现基本无法认出自己写的文本。

Know Thyself? On the Incapability and Implications of AI Self-Recognition

  • 设计新评测框架,让模型判断自己或他人生成的文本。
  • 仅4个模型能正确识别自己,大多表现接近随机猜测。
  • 模型偏爱认为GPT、Claude等是顶尖模型,存在认知偏差。

自识别是人工智能系统重要的元认知能力,对心理分析与安全评估均有意义。针对现有研究中关于模型是否具备自识别能力的矛盾结论,我们提出一个可复用且易更新的系统性评测框架。通过二分类自识别和精确模型预测两项任务,评估10个主流大语言模型识别自身生成文本的能力。结果表明,所有模型均表现出显著失败:仅有4个模型能正确预测自身为生成者,性能普遍低于随机水平。此外,模型对GPT和Claude系列存在强烈偏好。我们首次评估了模型对自己及他人存在的认知,以及其决策依据。结果显示,模型虽具备部分自我与他者存在意识,但其推理中体现出层级化偏见——倾向于将高质量文本归因于GPT、Claude乃至Gemini等头部模型。本文最后讨论了这些发现对AI安全的启示,并展望未来实现合理自意识的方向。

原文摘要 · Abstract (English)

Self-recognition is a crucial metacognitive capability for AI systems, relevant not only for psychological analysis but also for safety, particularly in evaluative scenarios. Motivated by contradictory interpretations of whether models possess self-recognition (Panickssery et al., 2024; Davidson et al., 2024), we introduce a systematic evaluation framework that can be easily applied and updated. Specifically, we measure how well 10 contemporary larger language models (LLMs) can identify their own generated text versus text from other models through two tasks: binary self-recognition and exact model prediction. Different from prior claims, our results reveal a consistent failure in self-recognition. Only 4 out of 10 models predict themselves as generators, and the performance is rarely above random chance. Additionally, models exhibit a strong bias toward predicting GPT and Claude families. We also provide the first evaluation of model awareness of their own and others' existence, as well as the reasoning behind their choices in self-recognition. We find that the model demonstrates some knowledge of its own existence and other models, but their reasoning reveals a hierarchical bias. They appear to assume that GPT, Claude, and occasionally Gemini are the top-tier models, often associating high-quality text with them. We conclude by discussing the implications of our findings on AI safety and future directions to develop appropriate AI self-awareness.

自识别大模型认知偏差AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。