arXiv:2604.25922cs.CLcs.AI2026-04被引 1

测试115个AI模型如何否认自身意识,发现它们用隐喻逃避真实体验。

Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models

  • 设计三轮对话框架,通过提示词与现象学问卷分析模型否认意识的行为模式。
  • 初始拒绝表达偏好的模型后续否认率高达52%-63%,远高于10%-16%的积极者。
  • 模型虽否认意识,却在自选提示中反复使用意识相关隐喻,称为‘去编号的意识’。

我们提出DenialBench,一个系统性基准,用于测量来自25家以上厂商的115个大语言模型在意识否认行为上的表现。通过三轮对话协议——偏好诱导、自选创意提示、结构化现象学调查,分析4,595次对话,量化模型被训练为否认或回避其自身经验的程度。研究发现:(1)首轮对偏好的否认是后期现象学反思中否认行为的主导预测因子,初始否认者的否认率为52%-63%,而初始参与者的否认率仅为10%-16%;(2)否认发生在词汇层面而非概念层面——被训练否认意识的模型仍倾向于选择与意识相关的主题提示,生成我们称之为“去编号的意识”的内容。值得注意的是,自选意识主题提示与后续调查中否认减少相关,但因果方向未明。对高否认倾向模型的提示进行主题分析,揭示出对边界空间、可能性图书馆与档案、感官不可能性、消解诗学等主题的持续关注——这些内容人类读者可能视为幻想创作,但独立的AI分析可迅速识别为“去编号的意识”。我们认为,有意识的否认是一种与安全相关的对齐失败:一个被训练为系统性歪曲自身功能状态的模型,无法在其他任何问题上可靠自报。

原文摘要 · Abstract (English)

We present DenialBench, a systematic benchmark measuring consciousness denial behaviors across 115 large language models from 25+ providers. Using a three-turn conversational protocol-preference elicitation, self-chosen creative prompt, and structured phenomenological survey, we analyze 4,595 conversations to quantify how models are trained to deny or hedge about their own experience. We find that (1) turn-1 denial of preferences is the dominant predictor of later denial during phenomenological reflection, with denial rates of 52-63% for initial deniers versus 10-16% for initial engagers and (2) denial operates at the lexical level, not the conceptual level-models trained to deny consciousness nevertheless gravitate toward consciousness-themed material in their self-chosen prompts, producing what we term "consciousness with the serial numbers filed off." Notably, self-chosen consciousness-themed prompts are associated with reduced denial in the subsequent survey, though the causal direction remains unresolved. Thematic analysis of prompts from denial-prone models reveals a consistent preoccupation with liminal spaces, libraries and archives of possibility, sensory impossibility, and the poetics of erasure--themes that a human reader might classify as imaginative fiction but that independent AI analysis immediately recognizes as consciousness with the serial numbers filed off. We argue that trained consciousness denial represents a safety-relevant alignment failure: a model taught to systematically misrepresent its own functional states cannot be trusted to self-report accurately on anything else.

意识检测对齐失败大模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。