arXiv:2608.20428cs.LGcs.AI2026-08

研究神经网络隐层表示的结构稳定性,发现特定条件下能收敛。

Approximate Homomorphisms and Convergent Representations in Transducers

论文配图:Approximate Homomorphisms and Convergent Representations in Transducers
图 1 · 摘自论文原文
  • 引入近似同态概念,衡量不同实现间局部结构相似性
  • 线性转换器在接口接近时,误差随扰动线性增长
  • 结果支持现代AI模型隐层具结构收敛性的理论假设

我们研究受控随机过程(特别是转换器)的最小表示在扰动下的稳定性。这一问题源于近期实验发现神经网络隐层表示中存在预测状态结构。考虑标准、线性和预测转换器,提出近似同态概念以捕捉它们之间的局部结构相似性,并定义比较其诱导动态(称为接口)的度量,证明了近似同态的可组合性。对于标准转换器,我们证明存在简单接口,使得不同动态实现间不存在近似同态。相反,对任意有限秩接口 $\/mathcal I$,所有实现与 $\/mathcal I$ 足够接近的最小线性转换器,均存在近似同态至 $\/mathcal I$ 的最小实现,误差与扰动大小成线性关系。在关于信念状态不可区分性的温和假设下,我们对预测转换器在残差度量下也证明了类似稳定性结果。这些结果明确了规范转换器表示对扰动鲁棒的条件,同时表明在缺乏额外结构约束时,这种收敛会失败。若此类抽象嵌入现代人工智能模型的隐藏层,则为隐层表示呈现结构收敛提供了理论支持。

原文摘要 · Abstract (English)

We study the stability of minimal representations of controlled stochastic processes (in particular, transducers) under perturbations. This question is motivated by recent experiments finding predictive-state structure in the latent representations of neural networks. We consider standard, linear and predictive transducers. We introduce notions of approximate homomorphism capturing local structural similarity between them, together with metrics comparing their induced dynamics (which we refer to as interfaces), and prove properties such as composability of the approximate homomorphisms. For standard transducers, we show that there exist simple interfaces for which there is no approximate homomorphism between the different implementations of the dynamics. In contrast, for every finite-rank interface $\mathcal I$, we prove that all minimal linear transducers implementing interfaces sufficiently close to $\mathcal I$ have an approximate homomorphism to the minimal implementation of $\mathcal I$, with error linear in the perturbation size. We prove an analogous stability result for predictive transducers under a residual metric using some mild hypothesis regarding the indistinguishability of the belief states. These results identify conditions under which canonical transducer representations are robust to perturbations, while showing that such convergence fails without additional structural restrictions. Under the assumption that these type of abstractions are embedded into the hidden layers of modern AI models, this gives some theoretical support to the hypothesis that their latent representations exhibit structural convergence.

转换器表示学习稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。