用混沌理论解释大模型推理机制,发现其依赖动态信息提取。
Cognitive Activation and Chaotic Dynamics in Large Language Models: A Quasi-Lyapunov Analysis of Reasoning Mechanisms
- 引入准李雅普诺夫指数量化模型各层混沌特性。
- 信息积累呈非线性指数增长,MLP贡献高于注意力机制。
- 微小初始扰动显著影响推理,适合研究模型可解释性者阅读。
大型语言模型(LLMs)表现出类人推理能力,挑战了传统神经网络理论对固定参数系统灵活性的理解。本文提出「认知激活」理论,从动力系统视角揭示了LLM推理机制的本质:模型的推理能力源于参数空间中动态信息提取的混沌过程。通过引入准李雅普诺夫指数(QLE),定量分析模型在不同层的混沌特征。实验表明,模型的信息积累遵循非线性指数规律,多层感知机(MLP)对最终输出的贡献高于注意力机制。进一步实验显示,微小初始值扰动会对模型推理能力产生显著影响,验证了大语言模型为混沌系统的理论分析。本研究为理解LLM推理的可解释性提供了混沌理论框架,并揭示了在模型设计中平衡创造力与可靠性的潜在路径。
原文摘要 · Abstract (English)
The human-like reasoning capabilities exhibited by Large Language Models (LLMs) challenge the traditional neural network theory's understanding of the flexibility of fixed-parameter systems. This paper proposes the "Cognitive Activation" theory, revealing the essence of LLMs' reasoning mechanisms from the perspective of dynamic systems: the model's reasoning ability stems from a chaotic process of dynamic information extraction in the parameter space. By introducing the Quasi-Lyapunov Exponent (QLE), we quantitatively analyze the chaotic characteristics of the model at different layers. Experiments show that the model's information accumulation follows a nonlinear exponential law, and the Multilayer Perceptron (MLP) accounts for a higher proportion in the final output than the attention mechanism. Further experiments indicate that minor initial value perturbations will have a substantial impact on the model's reasoning ability, confirming the theoretical analysis that large language models are chaotic systems. This research provides a chaos theory framework for the interpretability of LLMs' reasoning and reveals potential pathways for balancing creativity and reliability in model design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。