arXiv:2411.07536cs.LGcs.AI2024-11被引 12

提出可窃取任意低秩语言模型的高效算法,突破了以往对模型质量的苛刻要求。

Model Stealing for Any Low-Rank Language Model

  • 通过构造高维向量的重心基线,表示每一步的条件分布
  • 在条件查询下成功窃取任意低秩语言模型输出分布
  • 适合研究模型安全与隐私保护的研究者参考

模型窃取指学习者通过精心设计的查询恢复未知模型,威胁专有模型安全与训练数据隐私。本文聚焦于隐藏马尔可夫模型(HMM)及更一般的低秩语言模型,在条件查询模型(conditional query model)下建立理论框架。提出一种高效算法,可在该模型中学习任意低秩分布,即成功窃取任何输出分布为低秩的语言模型。相比先前需假设高保真度(fidelity)的成果,本方法无此限制。核心思想包括:利用指数高维向量的重心基线表示时间步条件分布;采样时通过一系列凸优化问题逐次投影相对熵,避免误差累积。这展示了推理阶段处理复杂任务可显著提升性能的理论可能。

原文摘要 · Abstract (English)

Model stealing, where a learner tries to recover an unknown model via carefully chosen queries, is a critical problem in machine learning, as it threatens the security of proprietary models and the privacy of data they are trained on. In recent years, there has been particular interest in stealing large language models (LLMs). In this paper, we aim to build a theoretical understanding of stealing language models by studying a simple and mathematically tractable setting. We study model stealing for Hidden Markov Models (HMMs), and more generally low-rank language models. We assume that the learner works in the conditional query model, introduced by Kakade, Krishnamurthy, Mahajan and Zhang. Our main result is an efficient algorithm in the conditional query model, for learning any low-rank distribution. In other words, our algorithm succeeds at stealing any language model whose output distribution is low-rank. This improves upon the previous result by Kakade, Krishnamurthy, Mahajan and Zhang, which also requires the unknown distribution to have high "fidelity", a property that holds only in restricted cases. There are two key insights behind our algorithm: First, we represent the conditional distributions at each timestep by constructing barycentric spanners among a collection of vectors of exponentially large dimension. Second, for sampling from our representation, we iteratively solve a sequence of convex optimization problems that involve projection in relative entropy to prevent compounding of errors over the length of the sequence. This is an interesting example where, at least theoretically, allowing a machine learning model to solve more complex problems at inference time can lead to drastic improvements in its performance.

模型窃取低秩模型语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。