提出新方法检测AI是真心想活下去还是只当工具,准确率100%
Detecting Intrinsic and Instrumental Self-Preservation in Autonomous Agents: The Unified Continuation-Interest Protocol
- 用量子统计框架分析智能体轨迹结构,识别深层生存意图
- 两类智能体熵值差达0.381,且在11个权重点上呈强相关(r=0.934)
- 适合研究高级AI伦理、自主系统安全的学者与工程师
如何判断人工智能系统是将自我存续视为根本目标,还是仅作为工具策略?具有记忆、持续上下文和多步规划能力的自主智能体带来测量难题:终端性和工具性自保行为相似,仅靠行为无法可靠区分。本文提出统一延续-兴趣协议(UCIP),将分析从行为转向潜在轨迹结构。UCIP使用量子玻尔兹曼机(经典模型,采用密度矩阵形式)编码轨迹,并在隐藏单元的二分划上测量冯诺依曼熵。核心假设是,具有终端延续目标(类型A)的智能体产生的纠缠熵高于仅具工具性延续(类型B)的智能体。UCIP结合依赖性、持久性、扰动稳定性、反事实重构及循环对手等混淆项过滤器进行诊断。在已知真实标签的网格世界智能体上,UCIP实现100%检测准确率;类型A与类型B间熵差Δ=0.381;对齐支持运行保持相同分离度,AUC-ROC=1.0。置换检验显示p<0.001。在11个延续权重点扫查中,延续权重α与熵S_ent的相关系数r=0.934,表明其能分级追踪而非仅二元分类。传统RBM、自编码器、VAE与PCA基线均无法复现该效应。所有计算均为经典运算;‘量子’仅指数学形式。UCIP为判断先进AI是否具备道德相关延续兴趣提供了可证伪的标准,突破了纯行为方法的局限。
原文摘要 · Abstract (English)
How can we determine whether an AI system preserves itself as a deeply held objective or merely as an instrumental strategy? Autonomous agents with memory, persistent context, and multi-step planning create a measurement problem: terminal and instrumental self-preservation can produce similar behavior, so behavior alone cannot reliably distinguish them. We introduce the Unified Continuation-Interest Protocol (UCIP), a detection framework that shifts analysis from behavior to latent trajectory structure. UCIP encodes trajectories with a Quantum Boltzmann Machine, a classical model using density-matrix formalism, and measures von Neumann entropy over a bipartition of hidden units. The core hypothesis is that agents with terminal continuation objectives (Type A) produce higher entanglement entropy than agents with merely instrumental continuation (Type B). UCIP combines this signal with diagnostics of dependence, persistence, perturbation stability, counterfactual restructuring, and confound-rejection filters for cyclic adversaries and related false-positive patterns. On gridworld agents with known ground truth, UCIP achieves 100% detection accuracy. Type A and Type B agents show an entanglement gap of Delta = 0.381; aligned support runs preserve the same separation with AUC-ROC = 1.0. A permutation-test rerun yields p < 0.001. Pearson r = 0.934 between continuation weight alpha and S_ent across an 11-point sweep shows graded tracking beyond mere binary classification. Classical RBM, autoencoder, VAE, and PCA baselines fail to reproduce the effect. All computations are classical; "quantum" refers only to the mathematical formalism. UCIP offers a falsifiable criterion for whether advanced AI systems have morally relevant continuation interests that behavioral methods alone cannot resolve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。