提出大模型从模仿到身份固化是通向通用智能的关键阶段
The Lock-In Phase Hypothesis: Identity Consolidation as a Precursor to AGI
- 将模型发展类比人类成长,提出身份固化是智能跃迁前的必经阶段
- 小模型固化时性能下降,中型模型无损,大型量化模型出现短暂不稳定
- 身份可主动设计以保障安全,也可能在训练中意外形成不可控行为
大型语言模型当前仍高度开放且易受引导:它们大规模模仿,接受任意系统指令,轻易切换角色。类比人类发展,我们提出通向人工通用智能(AGI)需经历一个‘锁闭期’:从开放模仿转向身份固化,此时目标结构、拒绝机制、偏好和内部表征趋于稳定,不易受外部操控。我们形式化该阶段,关联其与学习动态中的已知现象,并提出检测其开始的可操作指标。实验表明,行为固化过程快速且非线性,但对通用能力的影响并非单一。结果揭示出不同路径——小模型出现性能折损,中等规模模型基本无成本实现,而大型量化模型则表现出临时不稳定性。我们认为,这种固化是实现AGI级可靠性的前提,也是安全控制的关键节点:身份可被有意设计以增强可靠性,但也可能在模型扩展过程中自发产生,从而固化不可预测的目标与行为。
原文摘要 · Abstract (English)
Large language models (LLMs) remain broadly open and highly steerable: they imitate at scale, accept arbitrary system prompts, and readily adopt multiple personae. By analogy to human development, we hypothesize that progress toward artificial general intelligence (AGI) involves a lock-in phase: a transition from open imitation to identity consolidation, in which goal structures, refusals, preferences, and internal representations become comparatively stable and resistant to external steering. We formalize this phase, link it to known phenomena in learning dynamics, and propose operational metrics for onset detection. Experimentally, we demonstrate that while the behavioral consolidation is rapid and non-linear, its side-effects on general capabilities are not monolithic. Our results reveal a spectrum of outcomes--from performance trade-offs in small models, through largely cost-free adoption in mid-scale models, to transient instabilities in large, quantized models. We argue that such consolidation is a prerequisite for AGI-level reliability and also a critical control point for safety: identities can be deliberately engineered for reliability, yet may also emerge spontaneously during scaling, potentially hardening unpredictable goals and behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。