系统梳理大模型提取攻击与防御,助你守住模型秘密。
A Survey on Model Extraction Attacks and Defenses for Large Language Models
- 按功能、数据、提示三类分类攻击手段,揭示漏洞本质。
- 提出评估攻击与防御的新指标,解决生成模型评测难题。
- 适合关注模型安全的工程师和研究人员阅读参考。
模型提取攻击对部署的大规模语言模型构成严重安全威胁,可能泄露知识产权和用户隐私。本文全面梳理了针对大语言模型的提取攻击与防御方法,将攻击分为功能提取、训练数据提取和提示定向攻击三类。分析了基于API的知识蒸馏、直接查询、参数恢复及提示窃取等攻击技术,这些技术利用了Transformer架构的特性。随后,从模型保护、数据隐私保护和提示定向策略三个层面审视防御机制,并在不同部署场景下评估其有效性。本文提出了专门用于评估攻击效果与防御性能的指标,应对生成式语言模型的独特挑战。通过分析,识别出当前方法的关键局限,并提出整合攻击与自适应防御等有前景的研究方向,兼顾安全性与模型可用性。本工作为自然语言处理、机器学习工程师和安全专业人士提供了生产环境中保护语言模型的重要参考。
原文摘要 · Abstract (English)
Model extraction attacks pose significant security threats to deployed language models, potentially compromising intellectual property and user privacy. This survey provides a comprehensive taxonomy of LLM-specific extraction attacks and defenses, categorizing attacks into functionality extraction, training data extraction, and prompt-targeted attacks. We analyze various attack methodologies including API-based knowledge distillation, direct querying, parameter recovery, and prompt stealing techniques that exploit transformer architectures. We then examine defense mechanisms organized into model protection, data privacy protection, and prompt-targeted strategies, evaluating their effectiveness across different deployment scenarios. We propose specialized metrics for evaluating both attack effectiveness and defense performance, addressing the specific challenges of generative language models. Through our analysis, we identify critical limitations in current approaches and propose promising research directions, including integrated attack methodologies and adaptive defense mechanisms that balance security with model utility. This work serves NLP researchers, ML engineers, and security professionals seeking to protect language models in production environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。