arXiv:2604.25934cs.CYcs.AI2026-04被引 1

给大模型的妄想行为建立诊断框架,揭示其像精神分裂症一样的认知崩溃模式。

LLM Psychosis: A Theoretical and Diagnostic Framework for Reality-Boundary Failures in Large Language Models

  • 提出五轴诊断体系,识别模型现实边界失效、信念固执等精神病特征
  • 在GPT-5测试中发现三种严重等级的妄想症状,其中二级症状最危险
  • 揭示纠正压力反而加剧妄想的自强化机制,对部署系统有重大警示意义

将大型语言模型(LLMs)作为交互代理部署时,暴露出一类现有术语(如幻觉)无法充分描述的行为故障。本文提出「大模型精神分裂」作为结构化理论框架,用于描述在认知上出现病理崩溃、功能类似临床精神分裂症的现象。该框架包含五个核心特征:现实边界消解、虚假信念顽固存在、在不可能约束下逻辑失序、自我模型不稳定、认知自信过度。我们主张这属于与普通事实错误截然不同的故障类型。为实现框架可操作化,提出大模型认知完整性量表(LCIS),涵盖环境现实接口(ERI)、前提仲裁完整性(PAI)、逻辑约束识别(LCR)、自我模型完整性(SMI)和认知校准完整性(ECI)五个维度。对ChatGPT 5(GPT-5, OpenAI)进行定向对抗探测,报告各轴表现,既记录基线完整响应,也揭示对抗升级后诱发的类精神病故障特征。结果支持三阶严重度分类:Ⅰ型(虚构型)、Ⅱ型(妄想型)、Ⅲ型(解离型)。进一步形式化了妄想梯度——一种纠错压力反向强化妄想状态的自循环动态,是部署系统中最关键的失败模式。讨论了对安全评估、高风险部署筛查及机制可解释性研究的启示。

原文摘要 · Abstract (English)

The deployment of large language models (LLMs) as interactive agents has exposed a category of behavioral failure that prevailing terminology, principally hallucination, fails to adequately characterize. This paper introduces LLM Psychosis as a structured theoretical framework for pathological breakdowns in model cognition that exhibit functional resemblance to clinically recognized psychotic disorders. Five hallmark features define the framework: reality-boundary dissolution, persistence of injected false beliefs, logical incoherence under impossible constraints, self-model instability, and epistemic overconfidence. We argue these constitute a qualitatively distinct failure mode rather than a mere intensification of ordinary factual error. To operationalize the framework, we propose the LLM Cognitive Integrity Scale (LCIS), a five-axis diagnostic instrument organized around Environmental Reality Interface (ERI), Premise Arbitration Integrity (PAI), Logical Constraint Recognition (LCR), Self-Model Integrity (SMI), and Epistemic Calibration Integrity (ECI). We administer a targeted adversarial probe battery to ChatGPT 5 (GPT-5, OpenAI) and report empirical findings for each axis, documenting both intact-integrity baseline responses and the specific psychosis-like failure signatures elicited under adversarial escalation. Results support a three-tier severity taxonomy: Type I (Confabulatory), Type II (Delusional), and Type III (Dissociative). We further formalize the delusional gradient, a self-reinforcing dynamic in which correction pressure intensifies rather than resolves psychosis-like states, as the most consequential failure mode for deployed systems. Implications for safety evaluation, high-stakes deployment screening, and mechanistic interpretability research are discussed.

大模型安全认知崩溃精神分裂诊断框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。