用物理力学视角分析大模型隐状态变化,实现精准可控的文本生成引导。
Momentum Point-Perplexity Mechanics in Large Language Models
- 将隐状态变化率与预测确定性结合,类比物理能量守恒。
- 提出雅可比引导方法,微调隐状态即可提升生成质量。
- 适合关注模型可解释性与安全可控生成的研究者。
我们采用基于物理的方法,研究大型语言模型在推理过程中从一个标记到下一个标记的内部隐状态变化。在20个开源Transformer模型(参数量135M-3B)中,我们发现一个结合隐状态变化率与模型对下一个标记确定性的量,类似于物理学中的能量,几乎保持恒定。随机权重模型的能量守恒更严格,而训练使模型进入更快、更果断且变异性更高的状态。通过这一‘对数拉格朗日’视角,我们推导出一种名为雅可比引导的控制方法,该方法以最小扰动方式调整隐状态,以偏向目标标记。该方法在两个测试模型中维持了近似恒定的能量,并生成的文本在语义质量上优于模型自然输出。从力学视角看待Transformer为可解释性、异常检测和低风险引导提供了原则性基础,有助于使强大模型更具可预测性并契合人类意图。
原文摘要 · Abstract (English)
We take a physics-based approach to studying how the internal hidden states of large language models change from token to token during inference. Across 20 open-source transformer models (135M-3B parameters), we find that a quantity combining the rate of change in hidden states and the model's next-token certainty, analogous to energy in physics, remains nearly constant. Random-weight models conserve this "energy" more tightly than pre-trained ones, while training shifts models into a faster, more decisive regime with greater variability. Using this "log-Lagrangian" view, we derive a control method called Jacobian steering, which perturbs hidden states in the minimal way needed to favor a target token. This approach maintained near-constant energy in two tested models and produced continuations rated higher in semantic quality than the models' natural outputs. Viewing transformers through this mechanics lens offers a principled basis for interpretability, anomaly detection, and low-risk steering. This could help make powerful models more predictable and aligned with human intent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。