破解现代模型中无偏置GLU模块的密码学方法
Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation
- 通过反向查询与对称分离,重建孤立的GLU前馈块结构
- 六层Qwen模型、8192单元子问题等均实现亚百分比验证误差
- 适用于安全研究者分析模型内部组件,非完整攻击手段
已有密码学提取方法成功应用于ReLU网络、逐元素激活网络(如GELU或SiLU)以及Transformer的最终投影矩阵,但尚未能恢复现代语言模型中广泛使用的无偏置门控线性单元(GLU)前馈块。该结构在每个隐藏单元内将激活后的线性投影与另一学习得到的线性投影相乘,具有双分支特性,不同于前述方法所处理的网络类别和最终层设置。本文提出一种构造性的多阶段前向查询恢复原语,用于孤立的无偏置GLU块。利用有限差分曲率生成门控方向候选,通过在x和-x处的成对观测分离门控幅度、方向及值分支耦合。在高精度目标下,六层Qwen模型、8192单位的Llama子问题及全维度Gemma块均达到亚百分比中位数验证误差;四种有限精度配置中位误差低于5%,但均未完全复现所有存储权重。这些孤立块实验并非端到端模型接口攻击:从最终模型输出推导所需内部块响应仍为未解难题。
原文摘要 · Abstract (English)
Cryptanalytic extraction has been demonstrated for ReLU networks, for networks using componentwise activations such as GELU or SiLU, and for a Transformer's final projection matrix. These methods do not recover the bias-free Gated Linear Unit (GLU) feed-forward blocks used in many modern language models. Such a block multiplies an activated linear projection by a second learned linear projection within each hidden unit, a two-branch structure absent from the network classes and final-layer setting addressed by those methods. We give a constructive, multi-stage forward-query recovery primitive for isolated bias-free GLU blocks. Finite-difference curvature supplies gate-direction candidates, and paired observations at x and -x separate gate magnitude, orientation, and value-branch coupling. Across high-precision targets, six Qwen layers, an 8,192-unit Llama subproblem, and a full-dimensional Gemma block all reach sub-percent median validation error. Four finite-precision configurations remain below 5 percent median error, but none reproduces every stored weight. These isolated-block experiments are not an end-to-end model-API attack: deriving the required internal block responses from final model outputs remains unsolved.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。