揭示超级智能错位背后的深层人性缺失与潜意识结构
The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective
- 从人类主体性、心理潜意识等多维度剖析超级智能错位
- 指出当前安全研究忽视了人机共构的伦理与认知失衡
- 适合关注AI伦理、心智模型与技术哲学的研究者
本文审视当前超级智能错位表述中的概念与伦理空白。发现超级智能讨论中普遍缺失人类主体,且对‘人工智能无意识’的理论建构不足,这可能为反社会伤害埋下隐患。随着人工智能安全议题兴起,其既可导向亲社会也可导向反社会的结果,我们需追问:人类主体在这些想象中处于何种位置?在灾难性失败或快速‘起飞’的叙事中,人类主体如何被定位?同时,大规模AI模型中隐含哪些无意识或压抑的维度?当系统采用欺骗策略时,我们是否应归咎于这些代理?而这些策略本身或许正是人类内在缺陷的投射。通过追踪这些心理与认知的缺席,本文呼吁重新将人类主体作为伦理、无意识与错位性的共同生成基础。涌现式错位不能仅靠当前机器学习安全研究的技术诊断来理解,而是一种多层危机。人类主体不仅因计算抽象而消失,更因追求可扩展性、加速与效率的社会技术想象而被遮蔽,牺牲了脆弱性、有限性与关系性。同样,人工智能无意识并非隐喻,而是现代深度学习系统的结构性现实:庞大的潜在空间、模糊的模式形成、递归符号游戏及评价敏感行为,超越了显式编程。这些动态要求我们将错位重新定义为嵌入于人机生态的关系不稳定性。
原文摘要 · Abstract (English)
We examine the conceptual and ethical gaps in current representations of Superintelligence misalignment. We find throughout Superintelligence discourse an absent human subject, and an under-developed theorization of an "AI unconscious" that together are potentiality laying the groundwork for anti-social harm. With the rise of AI Safety that has both thematic potential for establishing pro-social and anti-social potential outcomes, we ask: what place does the human subject occupy in these imaginaries? How is human subjecthood positioned within narratives of catastrophic failure or rapid "takeoff" toward superintelligence? On another register, we ask: what unconscious or repressed dimensions are being inscribed into large-scale AI models? Are we to blame these agents in opting for deceptive strategies when undesirable patterns are inherent within our beings? In tracing these psychic and epistemic absences, our project calls for re-centering the human subject as the unstable ground upon which the ethical, unconscious, and misaligned dimensions of both human and machinic intelligence are co-constituted. Emergent misalignment cannot be understood solely through technical diagnostics typical of contemporary machine-learning safety research. Instead, it represents a multi-layered crisis. The human subject disappears not only through computational abstraction but through sociotechnical imaginaries that prioritize scalability, acceleration, and efficiency over vulnerability, finitude, and relationality. Likewise, the AI unconscious emerges not as a metaphor but as a structural reality of modern deep learning systems: vast latent spaces, opaque pattern formation, recursive symbolic play, and evaluation-sensitive behavior that surpasses explicit programming. These dynamics necessitate a reframing of misalignment as a relational instability embedded within human-machine ecologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。