arXiv:2602.03396cs.CL2026-02被引 3

从信息论角度提升大模型防知识蒸馏能力

Towards Distillation-Resistant Large Language Models: An Information-Theoretic Perspective

  • 用条件互信息量化教师输出中可被蒸馏的信息
  • 通过优化变换矩阵使输出去除了可提取信息但保持任务准确率
  • 适用于保护商业大模型知识产权的防御场景

专有的大语言模型蕴含巨大经济价值,通常仅以黑盒API形式提供,但攻击者仍可通过知识蒸馏获取其知识。现有防御主要针对文本蒸馏,忽略了重要的对数(logit)蒸馏。本文从信息论视角分析该问题,提出一种新方法:利用教师输出对数与输入查询在真实标签条件下的条件互信息(CMI),刻画可被蒸馏的信息。基于此,设计了一种输出变换矩阵,通过最小化CMI来净化输出,有效移除蒸馏相关特征。进一步提出基于CMI的抗蒸馏目标函数,在多个主流大模型和强蒸馏算法上验证表明,该方法显著降低蒸馏性能,同时维持原有任务准确性,有效保护模型知识产权。

原文摘要 · Abstract (English)

Proprietary large language models (LLMs) embody substantial economic value and are generally exposed only as black-box APIs, yet adversaries can still exploit their outputs to extract knowledge via distillation. Existing defenses focus exclusively on text-based distillation, leaving the important logit-based distillation largely unexplored. In this work, we analyze this problem and present an effective solution from an information-theoretic perspective. We characterize distillation-relevant information in teacher outputs using the conditional mutual information (CMI) between teacher logits and input queries conditioned on ground-truth labels. This quantity captures contextual information beneficial for model extraction, motivating us to defend distillation via CMI minimization. Guided by our theoretical analysis, we propose learning a transformation matrix that purifies the original outputs to enhance distillation resistance. We further derive a CMI-inspired anti-distillation objective to optimize this transformation, which effectively removes distillation-relevant information while preserving output utility. Extensive experiments across multiple LLMs and strong distillation algorithms demonstrate that the proposed method significantly degrades distillation performance while preserving task accuracy, effectively protecting models' intellectual property.

大模型安全知识蒸馏信息论模型防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。