arXiv:2601.08223cs.CRcs.AI2026-01中稿 · ICASSP2026被引 3

用双层嵌套指纹保护大模型版权,隐蔽且抗干扰。

DNF: Dual-Layer Nested Fingerprinting for Large Language Model Intellectual Property Protection

  • 结合风格特征与语义触发构建分层后门
  • 在多个模型上实现100%指纹激活且保持性能
  • 适合需要隐蔽确权的大模型开发者

大语言模型的快速普及带来了黑盒部署下的知识产权保护难题。现有基于后门的指纹技术要么依赖罕见词元,导致高困惑度输入易被过滤;要么使用固定触发-响应映射,易受泄露和后期适应攻击。本文提出双层嵌套指纹(DNF),通过耦合领域特异性风格线索与隐式语义触发,在 Mistral-7B、LLaMA-3-8B-Instruct 与 Falcon3-7B-Instruct 上实现完美指纹激活,同时保持下游任务性能。相比现有方法,其触发器困惑度更低,能抵御指纹检测攻击,对增量微调和模型融合也具备较强鲁棒性。DNF 是一种实用、隐蔽且强韧的 LLM 所有权验证与知识产权保护方案。

原文摘要 · Abstract (English)

The rapid growth of large language models raises pressing concerns about intellectual property protection under black-box deployment. Existing backdoor-based fingerprints either rely on rare tokens -- leading to high-perplexity inputs susceptible to filtering -- or use fixed trigger-response mappings that are brittle to leakage and post-hoc adaptation. We propose \textsc{Dual-Layer Nested Fingerprinting} (DNF), a black-box method that embeds a hierarchical backdoor by coupling domain-specific stylistic cues with implicit semantic triggers. Across Mistral-7B, LLaMA-3-8B-Instruct, and Falcon3-7B-Instruct, DNF achieves perfect fingerprint activation while preserving downstream utility. Compared with existing methods, it uses lower-perplexity triggers, remains undetectable under fingerprint detection attacks, and is relatively robust to incremental fine-tuning and model merging. These results position DNF as a practical, stealthy, and resilient solution for LLM ownership verification and intellectual property protection.

模型确权后门指纹大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。