arXiv:2509.03122cs.CLcs.AI2025-09被引 2

用代码混杂与多重编辑技术,实现难以察觉且抗修改的模型指纹。

From Construction to Injection: Edit-Based Fingerprints for Large Language Models

  • 通过低困惑度代码混杂,在隐蔽性与安全性间取得平衡。
  • 多候选编辑构建冗余触发机制,支持模型修改后的持续验证。
  • 适用于防范模型盗用,适合关注版权保护的研究者。

可靠的模型指纹对防止大语言模型被非法分发和商业滥用至关重要。在黑盒部署中,验证因可疑指纹查询的防御性过滤以及下游模型修改导致的版权证据弱化而受阻。因此,指纹需在构建和注入两方面均具备鲁棒性。现有方法在构建上面临不可察觉性权衡:自然语言指纹可能误激活,而混乱指纹则易被统计识别并过滤。在注入上,现有方法难以在模型修改后保持持久的触发-目标行为。本文提出端到端注入指纹框架,采用最低困惑度代码混杂(CF)在高复杂度约束下缓解双重不可察觉性矛盾;通过多候选编辑(MCEdit)构建结构冗余、间隔分离的触发-目标映射,实现模型修改下的渐进式退化。大量实验表明,该方法在不可察觉性、可检测性和无害性方面表现优异,对模型性能影响极小。

原文摘要 · Abstract (English)

Reliable model fingerprints are essential for protecting large language models (LLMs) against unauthorized redistribution and commercial misuse. In black-box deployment, verification is hindered by defensive filtering of suspected fingerprint queries, as well as by downstream model modifications that may weaken embedded ownership evidence. These risks require fingerprints to be robust in both construction and injection. For construction, prior paradigms face an imperceptibility trade-off: natural-language fingerprints may be accidentally activated, whereas garbled fingerprints are statistically exposed and easier to filter. For injection, existing methods struggle to preserve persistent trigger--target behaviors under model modification. We propose an end-to-end injected fingerprinting framework to address these challenges. Code-mixing Fingerprints (CF) use lowest-perplexity code-mixing under a high-complexity constraint to mitigate this two-sided imperceptibility trade-off. Multi-Candidate Editing (MCEdit) constructs structurally redundant, margin-separated trigger--target mappings to enable graceful degradation under model modification. Extensive evaluations on imperceptibility, detectability, and harmlessness demonstrate robust ownership verification with negligible impact on utility.

模型指纹版权保护安全鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。