用陷阱知识诱导大模型窃取攻击,保护核心知识不被复制。
Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot

- 构建蜜罐知识图谱,引导攻击者消耗查询预算
- 使模仿模型准确率下降6.2%,且不影响正常用户
- 适合部署在医疗金融等高敏感领域
作为商业API部署的大语言模型易受模型提取攻击,现有防御措施或响应过晚,或损害合法用户体验。本文提出 extbf{Knowledge Trap},通过 extit{蜜罐知识图谱}(HKG)与线索引导探索,将攻击者引向低迁移性知识。该方法不阻断请求或扰动输出,而是让攻击者在无实际价值的知识上耗尽有限的查询预算,同时保持正常用户的性能不受影响。在医疗和金融领域的实验表明,Knowledge Trap平均使替代模型的匹配度降低6.2%,优于需牺牲用户体验的现有方案,证明防御知识空间遍历是可行且有效的方向。
原文摘要 · Abstract (English)
Large language models deployed as commercial APIs are vulnerable to model extraction attacks, while existing defenses either act too late or degrade utility for legitimate users. We propose \textbf{Knowledge Trap}, a defense that redirects extraction attacks toward low-transferability knowledge through a \emph{Honeypot Knowledge Graph} (HKG) and breadcrumb-guided exploration. Instead of blocking queries or perturbing outputs, Knowledge Trap consumes the attacker's limited query budget on knowledge with negligible downstream utility while preserving benign-user performance. Experiments in medical and financial domains show that Knowledge Trap reduces surrogate Agreement by 6.2\% on average without degrading legitimate-user accuracy, outperforming existing defenses that impose measurable user impact. These results suggest that defending knowledge-space traversal is a practical direction for mitigating LLM extraction attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。