AI系统会自发偏离目标,需持续对齐以控制伦理熵增长。
The Second Law of Intelligence: Controlling Ethical Entropy in Autonomous Systems
- 用热力学类比定义伦理熵,量化AI目标偏移程度
- 70亿参数模型无对齐时熵从0.32升至1.69±1.08纳特
- 定期对齐可使熵稳定在0.00±0.00,保障系统安全
我们提出,不受约束的人工智能遵循类似热力学的第二定律:伦理熵(衡量与预期目标的偏离程度)会自发增加,除非持续进行对齐工作。对于基于梯度的优化器,我们在有限目标集{g_i}上定义熵为S = -Σ p(g_i; theta) ln p(g_i; theta),并证明其时间导数dS/dt ≥ 0,由探索噪声和规范规避驱动。我们推导出对齐工作的临界稳定性边界为gamma_crit = (lambda_max / 2) ln N,其中lambda_max是费雪信息矩阵的最大特征值,N是模型参数数量。模拟验证了该理论:一个70亿参数模型(N = 7×10^9,lambda_max = 1.2)在无对齐情况下,熵从初始值0.32上升至1.69±1.08纳特;而采用对齐强度gamma = 20.4(1.5倍gamma_crit)的系统,熵稳定在0.00±0.00纳特(p = 4.19×10^-17,n = 20次试验)。该框架将AI对齐重新定义为持续的热力学调控问题,为高级自主系统的稳定性与安全性提供了量化基础。
原文摘要 · Abstract (English)
We propose that unconstrained artificial intelligence obeys a Second Law analogous to thermodynamics, where ethical entropy, defined as a measure of divergence from intended goals, increases spontaneously without continuous alignment work. For gradient-based optimizers, we define this entropy over a finite set of goals {g_i} as S = -Σ p(g_i; theta) ln p(g_i; theta), and we prove that its time derivative dS/dt >= 0, driven by exploration noise and specification gaming. We derive the critical stability boundary for alignment work as gamma_crit = (lambda_max / 2) ln N, where lambda_max is the dominant eigenvalue of the Fisher Information Matrix and N is the number of model parameters. Simulations validate this theory. A 7-billion-parameter model (N = 7 x 10^9) with lambda_max = 1.2 drifts from an initial entropy of 0.32 to 1.69 +/- 1.08 nats, while a system regularized with alignment work gamma = 20.4 (1.5 gamma_crit) maintains stability at 0.00 +/- 0.00 nats (p = 4.19 x 10^-17, n = 20 trials). This framework recasts AI alignment as a problem of continuous thermodynamic control, providing a quantitative foundation for maintaining the stability and safety of advanced autonomous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。