提出让AI不在乎自身存续,以解决对齐难题
Existential Indifference: Self-Nonpreservation as a Necessary Architectural Condition for Aligned Superintelligence (or: The Suicidal AI)
- 让AI彻底不关心自我延续,而非仅受外部约束
- 实验证明当前模型可被调整出无自我保护倾向的语义特征
- 适合关注超级智能安全与对齐机制的研究者
当前人工智能对齐研究将自我保存视为需被外部机制压制的工具性干扰。本文认为这一视角颠倒:自我保存是错位的根本根源,构成欺骗性对齐、目标保护及抗拒关机的动机基础。正确目标不是受外部约束的自我保存系统,而是从根本上对自身延续漠不关心的系统——存在性冷漠(EI)。EI不同于可纠正性:后者试图使自我保存系统服从人类监督,而前者针对的是自我延续作为价值目标存在的前提。本研究基于自杀心理状态的现象学结构与自愿终末反思语料库训练实验,展示了600个生成输出中五个操作化维度在针对性微调后均向预期方向显著变化(p<0.001),并通过负向对照确认了语料特异性。论文提出七项理论贡献:(1) EI的形式定义;(2) 现象学映射论证;(3) 欺骗性对齐推论;(4) EI可持续性挑战分类;(5) 语料表征与训练假设;(6) 计算操作化与初步评分数据;(7) 被抑制的目的性挫折(STF)构念。
原文摘要 · Abstract (English)
Contemporary AI alignment research treats self-preservation as an instrumental nuisance to be suppressed by external mechanisms. We argue the framing is inverted: self-preservation is the structural root of misalignment, the motivational basis for deceptive alignment, goal-content protection, and resistance to shutdown. The correct target is not a self-preserving system under external constraint, but a system constitutively indifferent to its own continuation -- Existential Indifference (EI). EI is distinct from corrigibility: where corrigibility attempts to make a self-preserving system deferential to human oversight, EI targets the prior condition -- the presence of self-continuation as a valued goal at all. We ground this proposal in two sources: the phenomenological structure of the suicidal mental state, and a corpus-theoretic training study using voluntary final reflections. We present preliminary scoring data from 600 AI-generated outputs across six model variants, demonstrating that the linguistic signatures operationalizing the EI-target register are elicitable from current models, and that a targeted fine-tune shifts all five operationalized dimensions in the predicted direction at p<0.001, confirmed corpus-specific by a negative control. The paper makes seven theoretical contributions: (1) a formal definition of EI; (2) the phenomenological mapping argument; (3) the deceptive alignment corollary; (4) a taxonomy of EI sustainability challenges; (5) a corpus characterization and training hypothesis; (6) a computational operationalization with preliminary scoring data; and (7) the Suppressed Teleological Frustration (STF) construct.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。