用熵量化大模型价值漂移,实现实时安全监控
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
- 引入伦理熵概念,通过行为分类器动态评估模型价值观变化
- 指令微调使伦理熵降低约80%,基线模型熵持续上升
- 构建监控系统,实时预警价值漂移,适合模型部署安全团队
大型语言模型的安全性通常通过静态基准测试评估,但关键风险具有动态性:分布偏移下的价值漂移、越狱攻击以及部署中的对齐退化。基于近期提出的智能第二定律——将伦理熵视为状态变量,除非被对齐工作抵消,否则趋向增加——我们使该框架在大语言模型中可操作化。定义五类行为分类体系,训练分类器从模型输出中估计伦理熵 S(t),并在压力测试中测量四个前沿模型的基线与指令微调版本的熵动态。基线模型表现出持续的熵增长,而微调版本抑制了漂移,使伦理熵降低约80%。基于这些轨迹,我们估算有效对齐工作率 gamma_eff,将 S(t) 与 gamma_eff 嵌入监控管道,当熵漂移超过稳定性阈值时触发警报,实现对价值漂移的运行时监督。
原文摘要 · Abstract (English)
Large language model safety is usually assessed with static benchmarks, but key failures are dynamic: value drift under distribution shift, jailbreak attacks, and slow degradation of alignment in deployment. Building on a recent Second Law of Intelligence that treats ethical entropy as a state variable which tends to increase unless countered by alignment work, we make this framework operational for large language models. We define a five-way behavioral taxonomy, train a classifier to estimate ethical entropy S(t) from model transcripts, and measure entropy dynamics for base and instruction-tuned variants of four frontier models across stress tests. Base models show sustained entropy growth, while tuned variants suppress drift and reduce ethical entropy by roughly eighty percent. From these trajectories we estimate an effective alignment work rate gamma_eff and embed S(t) and gamma_eff in a monitoring pipeline that raises alerts when entropy drift exceeds a stability threshold, enabling run-time oversight of value drift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。