揭示Transformer隐藏状态的唯一性与几何鲁棒性,为模型可逆性提供理论和实证支持。
Transformer Injectivity & Geometric Robustness - Analytic Margins and Bi-Lipschitz Uniformity of Sequence-Level Hidden States
- 从解析角度证明了Transformer在有限提示集上隐状态映射的普遍单射性。
- 实测显示全精度与8位量化下无碰撞,4位量化导致少量碰撞且鲁棒性下降。
- 提出几何诊断工具,适用于分析模型规模、层深及量化对表示质量的影响。
在解码器仅型Transformer的实解析假设下,近期工作表明从离散提示到最后一层隐藏状态的映射在有限提示集上通常是单射的。本文进一步细化该结果:对每一层ℓ,定义碰撞判别集Δℓ ⊂ Θ 和单射层流Uℓ = Θackslash Δℓ,证明二分定理——要么模型在该集合上处处非单射,要么Uℓ是开稠密集且每个Fℓ_θ均为单射。在优化器非奇异性和绝对连续初始化的温和假设下,单射性沿任意固定时间范围的平滑训练轨迹保持不变。同时考虑对称群G,证明判别集和单射层流可降维至商空间Θ/G,因此单射性天然属于函数等价类。通过大规模提示集上的最近邻统计,我们定义了提示空间与最后一层表示空间之间的分离边界与共利普希茨(下利普希茨)常数,并应用于预训练的LLaMA-3和Qwen模型,研究不同层、序列长度、模型尺度以及8位和4位激活量化下的行为。在采样提示中,全精度与8位量化均未观察到碰撞,而4位量化引入少量碰撞并显著降低共利普希茨估计值。对从头训练的小型GPT-2,归一化指标在训练过程中保持稳定。总体表明,在连续参数理想化下,Transformer表示具有普遍且持久的单射性,其实际可逆性可通过简单的几何诊断工具探测。
原文摘要 · Abstract (English)
Under real-analytic assumptions on decoder-only Transformers, recent work shows that the map from discrete prompts to last-token hidden states is generically injective on finite prompt sets. We refine this picture: for each layer $\ell$ we define a collision discriminant $Δ^\ell \subset Θ$ and injective stratum $U^\ell = Θ\setminus Δ^\ell$, and prove a dichotomy -- either the model is nowhere injective on the set, or $U^\ell$ is open and dense and every $F^\ell_θ$ is injective. Under mild non-singularity assumptions on the optimizer and an absolutely continuous initialization, generic injectivity persists along smooth training trajectories over any fixed horizon. We also treat symmetry groups $G$, showing that discriminants and injective strata descend to the quotient $Θ/G$, so injectivity is naturally a property of functional equivalence classes. We complement these results with an empirical study of layerwise geometric diagnostics. We define a separation margin and a co-Lipschitz (lower Lipschitz) constant between prompt space and last-token representation space, estimated via nearest-neighbor statistics on large prompt sets. Applying these diagnostics to pretrained LLaMA-3 and Qwen models, we study behavior across layers, sequence lengths, model scales, and 8- and 4-bit activation quantization. On our sampled prompts we see no collisions in full precision or at 8 bits, while 4-bit quantization induces a small number of collisions and markedly shrinks co-Lipschitz estimates. For a small GPT-2 trained from scratch, normalized metrics remain stable over training. Overall, the results suggest that Transformer representations are generically and persistently injective in the continuous-parameter idealization, while their practical invertibility can be probed using simple geometric diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。