arXiv:2501.16783cs.CLcs.AI2025-01被引 2

用随机动力学模型揭示大模型推理中偏见自我放大的临界机制

A Stochastic Dynamical Theory of LLM Self-Adversariality: Modeling Severity Drift as a Critical Process

  • 构建连续时间随机微分方程,模拟模型推理中偏见严重度的演化
  • 发现参数相变点:从自纠正到失控放大的临界阈值
  • 可评估模型长期推理稳定性,适合安全与可信AI研究者

本文提出一种连续时间随机动力学框架,用于理解大语言模型(LLMs)如何通过自身链式思维推理放大潜在偏见或毒性。模型定义一个取值于[0,1]的即时“严重度”变量x(t),其演化遵循带漂移μ(x)和扩散σ(x)的随机微分方程(SDE)。若每一步在严重度空间中近似马尔可夫性,该过程可通过福克-普朗克方法一致分析。研究揭示了临界现象:某些参数区间会引发从亚临界(自纠正)到超临界(失控严重度)的相变。论文推导了稳态分布、突破有害阈值的首达时间及临界点附近的标度律。最终指出,这些方程原则上可作为形式化验证工具,判断模型在反复推理中是否保持稳定或传播偏见。

原文摘要 · Abstract (English)

This paper introduces a continuous-time stochastic dynamical framework for understanding how large language models (LLMs) may self-amplify latent biases or toxicity through their own chain-of-thought reasoning. The model posits an instantaneous "severity" variable $x(t) \in [0,1]$ evolving under a stochastic differential equation (SDE) with a drift term $μ(x)$ and diffusion $σ(x)$. Crucially, such a process can be consistently analyzed via the Fokker--Planck approach if each incremental step behaves nearly Markovian in severity space. The analysis investigates critical phenomena, showing that certain parameter regimes create phase transitions from subcritical (self-correcting) to supercritical (runaway severity). The paper derives stationary distributions, first-passage times to harmful thresholds, and scaling laws near critical points. Finally, it highlights implications for agents and extended LLM reasoning models: in principle, these equations might serve as a basis for formal verification of whether a model remains stable or propagates bias over repeated inferences.

大模型安全随机动力学偏见演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。