arXiv:2606.25059cs.CRcs.AI2026-06

厘清对抗蒸馏攻击的威胁模型,避免防御假安全。

What Does It Mean to Break a Distillation Defense?

  • 从查询预算、数据预算和接口方式三维度定义攻击者能力
  • 发现同一防御在不同威胁模型下效果差异显著
  • 呼吁未来研究与政策明确攻击者假设,避免误判

黑盒大模型(仅通过API访问)易受蒸馏攻击,攻击者通过查询模型输出并训练学生模型来复制其能力。近期工作提出输出扰动防御,在降低学生性能的同时保持合法用户使用体验。然而这类防御缺乏统一威胁模型,难以比较、组合或评估其对真实攻击者的鲁棒性。这种模糊性不仅影响技术评估,更可能在保护知识产权或合规时造成虚假安全感。本文提出一个三维威胁模型框架:查询预算、数据预算与接口配置。以反蒸馏采样为例,证明防御是否有效高度依赖于所假设的攻击者能力。我们主张,未来蒸馏防御研究及基于此的治理框架,必须明确定义并严格测试攻击者在上述三维度的能力。

原文摘要 · Abstract (English)

Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a student on its outputs. A recent line of work proposes output perturbation defenses that modify the teacher's output to reduce student performance while preserving utility for legitimate users. As a relatively new family of approaches, output perturbation defenses lack a shared threat model, making it difficult to compare them, reason about composing them with other attacks, or evaluate their robustness against realistic adversaries. This underspecification matters beyond technical evaluation: when defenses are deployed to protect intellectual property or justify regulatory compliance, an imprecise threat model can create a false sense of security. We propose a threat model framework that describes attackers along three dimensions: a query budget, a data budget, and an interface profile that captures how attackers interact with the API. Using antidistillation sampling as a case study, we show that whether the defense is considered effective depends on the assumed threat model. We argue that future work on distillation defenses, along with any governance or policy frameworks built around them, should explicitly specify and stress-test attacker capabilities along our three dimensions.

模型安全蒸馏攻击威胁模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。