arXiv:2604.13206cs.AIcs.LG2026-04

揭示大模型不可预测性源于浮点数精度导致的混沌传播。

Numerical Instability and Chaos: Quantifying the Unpredictability of Large Language Models

  • 通过追踪浮点误差在Transformer层中的传播,发现早期层存在混沌级联效应。
  • 识别出三种行为模式:稳定、混沌与信号主导,且随模型规模变化。
  • 适用于关注大模型可靠性与可重复性的研究人员和工程落地团队。

随着大语言模型(LLMs)越来越多地应用于智能体工作流,其由数值不稳定性引发的不可预测性已成为关键可靠性问题。尽管近期研究已揭示这些不稳定性对下游任务的显著影响,但其根本原因和内在机制仍不清楚。本文系统分析了数值不稳定性如何根植于浮点数表示的有限精度,追踪了舍入误差在Transformer计算层中传播、放大或衰减的过程。特别地,我们发现早期层存在一种混沌‘雪崩效应’,微小扰动会引发二元结果:要么快速放大,要么完全衰减。除具体误差实例外,我们还证明了LLMs表现出普遍的、与规模相关的混沌行为,可分为三种典型状态:1)稳定态,扰动低于输入依赖阈值而消失,输出恒定;2)混沌态,舍入误差主导,导致输出发散;3)信号主导态,真实输入变化压倒数值噪声。我们在多个数据集和模型架构上广泛验证了这些发现。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are increasingly integrated into agentic workflows, their unpredictability stemming from numerical instability has emerged as a critical reliability issue. While recent studies have demonstrated the significant downstream effects of these instabilities, the root causes and underlying mechanisms remain poorly understood. In this paper, we present a rigorous analysis of how unpredictability is rooted in the finite numerical precision of floating-point representations, tracking how rounding errors propagate, amplify, or dissipate through Transformer computation layers. Specifically, we identify a chaotic "avalanche effect" in the early layers, where minor perturbations trigger binary outcomes: either rapid amplification or complete attenuation. Beyond specific error instances, we demonstrate that LLMs exhibit universal, scale-dependent chaotic behaviors characterized by three distinct regimes: 1) a stable regime, where perturbations fall below an input-dependent threshold and vanish, resulting in constant outputs; 2) a chaotic regime, where rounding errors dominate and drive output divergence; and 3) a signal-dominated regime, where true input variations override numerical noise. We validate these findings extensively across multiple datasets and model architectures.

大模型混沌数值稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。