arXiv:2607.28667cs.LGcs.AI2026-07被引 1

用动力系统理论解释大模型生成文本的可区分性

Guarantees on Dynamical System Distinguishability for LLM Token Generation

论文配图:Guarantees on Dynamical System Distinguishability for LLM Token Generation
图 1 · 摘自论文原文
  • 将文本生成建模为两个随机线性动力系统的二元假设检验
  • 分类错误率随序列长度指数下降,由动力系统谱距决定
  • 揭示了跨嵌入模型泛化能力的数学边界,适合理论研究者

近期研究表明,可通过将语言模型的标记嵌入视为黑箱动力系统(DS)的轨迹,并比较两个DS的预测残差来区分其输出。尽管该方法在实践中表现良好,但对其为何有效、性能如何随序列长度变化以及能否跨嵌入模型迁移仍缺乏理论理解。本文将分类任务形式化为两个随机线性动力系统的二元假设检验。结果表明,即使动力学差异显著,两个系统的平稳边缘分布之间的总变差距离仍可任意小,这为忽略动态信息的分类器设定了基本准确率下限。进一步证明,基于动力系统的分类错误率随序列长度 $L$ 指数衰减,衰减速率由衡量两系统谱距的动力可区分性量 $δ^2$ 决定。同时通过引入嵌入模型间的近似纠缠条件,刻画了跨模型泛化能力,并建立了转移可区分性的下界,该下界与纠缠映射的最小奇异值相关。这些结果共同解释了动力系统方法的实证表现,并推动以动力系统理论分析AI系统,而非传统上用AI建模动力系统。

原文摘要 · Abstract (English)

Recent work has shown that classifying large language models (LLMs)' responses can be distinguished by modeling token embeddings as trajectories of a black-box dynamical system (DS) and comparing prediction residuals of two DSs. Despite the empirical success of this dynamical approach, a theoretical understanding of why it works, how well it scales as a function of the token sequence, and when it transfers across embedding models remains lacking. We address these questions by formalizing the classification task as a binary hypothesis test between two stochastic linear DSs. We show that the total variation distance between the stationary marginal distributions of the two DSs can be arbitrarily small even when the dynamics differ substantially, which provides a fundamental accuracy floor for any classifier that ignores token dynamics. We then show that the misclassification probability of DS-based classification decays exponentially in the sequence length $L$, with the decay governed by a dynamical discriminability quantity $δ^2$ that captures the spectral distance between the two DSs. We also characterize cross-embedding generalization by introducing an approximate intertwining condition between embedding models and establishing a lower bound on the transferable discriminability in terms of the intertwining map's smallest singular value. Together, these results explain the empirical performance of DS-based classification and motivate further investigation into using DS theory to analyze AI systems, in contrast to the more common approach of using AI to model dynamical systems.

大模型分析动力系统可区分性理论保障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。