arXiv:2510.16968cs.LGcs.AI2025-10

通过专家路由模式识别知识蒸馏,有效防避提示欺骗。

Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures

  • 利用MoE模型的专家分工与路由模式作为蒸馏指纹。
  • 在多种场景下检测准确率超94%,且抗提示攻击能力强。
  • 适用于任意模型对,支持白盒与黑盒检测,适合模型版权保护研究者。

知识蒸馏(KD)虽能加速大语言模型(LLMs)训练,但带来知识产权保护与模型多样性风险。现有基于自身份辨或输出相似性的检测方法易被提示工程规避。本文提出一种新框架,通过挖掘被忽视的信号——MoE的“结构习惯”,尤其是内部路由模式,在白盒和黑盒设置下实现高效检测。该方法分析不同专家在各类输入下的专业化与协作方式,形成持久的特征指纹。为拓展至非MoE架构与黑盒场景,进一步提出Shadow-MoE:通过辅助蒸馏构建代理MoE表示,比较任意模型对间的路由模式差异。建立了一个可复现的全面基准,包含多样化的蒸馏检查点与可扩展框架。大量实验表明,该方法在多种场景下检测准确率超过94%,对提示攻击具有强鲁棒性,显著优于现有基线,并揭示了大型语言模型中结构习惯的迁移现象。

原文摘要 · Abstract (English)

Knowledge Distillation (KD) accelerates training of large language models (LLMs) but poses intellectual property protection and LLM diversity risks. Existing KD detection methods based on self-identity or output similarity can be easily evaded through prompt engineering. We present a KD detection framework effective in both white-box and black-box settings by exploiting an overlooked signal: the transfer of MoE "structural habits", especially internal routing patterns. Our approach analyzes how different experts specialize and collaborate across various inputs, creating distinctive fingerprints that persist through the distillation process. To extend beyond the white-box setup and MoE architectures, we further propose Shadow-MoE, a black-box method that constructs proxy MoE representations via auxiliary distillation to compare these patterns between arbitrary model pairs. We establish a comprehensive, reproducible benchmark that offers diverse distilled checkpoints and an extensible framework to facilitate future research. Extensive experiments demonstrate >94% detection accuracy across various scenarios and strong robustness to prompt-based evasion, outperforming existing baselines while highlighting the structural habits transfer in LLMs.

知识蒸馏模型检测MoE安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。