arXiv:2606.06854cs.LG2026-06

用几何方法揭示如何精准窃取模型最后一层,明确边界。

The Geometry of Last-Layer Model Stealing

  • 基于几何分析,给出完美复制Transformer最后一层的条件。
  • 发现隐藏层无法仅凭输出结果完全逆向重构。
  • 为模型安全提供清晰的可盗与不可盗边界,适合安全研究者参考。

本文通过几何视角解释一种已知的机器学习模型窃取方法。作者精确推导出完全复制Transformer网络最后一层所需的条件。深入分析隐藏层后,明确了其重构的理论极限。研究表明,仅通过观察最终输出结果,无法完全逆向工程隐藏网络结构。该研究系统地划清了模型可被窃取与不可被窃取的边界。

原文摘要 · Abstract (English)

This paper uses geometry to explain how a machine learning model can be stolen using an already existing well-known method. The author has shown the exact conditions required to perfectly copy the final layer of a transformer network. When looking deeper into the hidden layers the author has explained clear limits. The author has also demonstrated that a hidden network cannot be fully reverse engineered just by looking at the final results. The research clearly maps out what can and cannot be stolen from a model.

模型安全几何分析模型窃取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。