arXiv:2508.20211cs.LGcs.SY2025-08

将Transformer视为非线性预测器,揭示其与经典滤波理论的联系。

What can we learn from signals and systems in a transformer? Insights for probabilistic modeling and inference architecture

  • 用概率模型解释Transformer信号为条件分布的近似
  • 层操作被建模为固定点迭代更新
  • 在隐马尔可夫模型下给出显式更新公式

20世纪40年代,维纳提出线性预测器,未来预测通过线性组合历史数据得到。Transformer则推广此思想:它是一个非线性预测器,下一个词的预测由非线性组合历史词元生成。本文提出一个概率模型,将Transformer中的信号视为条件测度的代理,将层操作视为固定点更新。当概率模型为隐马尔可夫模型(HMM)时,给出了固定点更新的显式形式。本工作部分旨在连接经典非线性滤波理论与现代推理架构。

原文摘要 · Abstract (English)

In the 1940s, Wiener introduced a linear predictor, where the future prediction is computed by linearly combining the past data. A transformer generalizes this idea: it is a nonlinear predictor where the next-token prediction is computed by nonlinearly combining the past tokens. In this essay, we present a probabilistic model that interprets transformer signals as surrogates of conditional measures, and layer operations as fixed-point updates. An explicit form of the fixed-point update is described for the special case when the probabilistic model is a hidden Markov model (HMM). In part, this paper is in an attempt to bridge the classical nonlinear filtering theory with modern inference architectures.

Transformer概率建模滤波理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。