提出隐式指纹技术,让大模型版权信息更隐蔽安全。
ImF: Implicit Fingerprint for Large Language Models
- 用隐写术将版权信息嵌入自然文本中
- 在15个不同模型上验证,指纹几乎无法被攻击移除
- 适合关注模型版权保护的研究者和开发者
训练大语言模型(LLMs)成本高昂,保护其知识产权至关重要。现有指纹嵌入方法通常在输出中加入语义不连贯的可识别模式,导致指纹与模型固有的问答行为存在显著差异,降低隐蔽性并易受对抗攻击。本文首次通过提出新型对抗攻击——生成重写干预(GRI)攻击,揭示了现有方法的严重漏洞:该攻击利用指纹语义结构脆弱性,有效擦除指纹。实证评估显示,传统方法在真实对抗环境下显著失效。为此,我们提出新范式隐式指纹(ImF),采用隐写术将所有权信息嵌入自然文本,并结合思维链(CoT)提示生成语义连贯、上下文自然的问答对,使指纹与模型正常输出难以区分,大幅降低被意外触发或针对性移除的风险。我们在15个涵盖不同架构与规模的LLM上进行了全面评估。
原文摘要 · Abstract (English)
Training large language models (LLMs) is resource-intensive and expensive, making protecting intellectual property (IP) for LLMs crucial. Recently, embedding fingerprints into LLMs has emerged as a prevalent method for establishing model ownership. However, existing fingerprinting techniques typically embed identifiable patterns with weak semantic coherence, resulting in fingerprints that significantly differ from the natural question-answering (QA) behavior inherent to LLMs. This discrepancy undermines the stealthiness of the embedded fingerprints and makes them vulnerable to adversarial attacks. In this paper, we first demonstrate the critical vulnerability of existing fingerprint embedding methods by introducing a novel adversarial attack named Generation Revision Intervention (GRI) attack. GRI attack exploits the semantic fragility of current fingerprinting methods, effectively erasing fingerprints by disrupting their weakly correlated semantic structures. Our empirical evaluation highlights that traditional fingerprinting approaches are significantly compromised by the GRI attack, revealing severe limitations in their robustness under realistic adversarial conditions. To advance the state-of-the-art in model fingerprinting, we propose a novel model fingerprint paradigm called Implicit Fingerprints (ImF). ImF leverages steganography techniques to subtly embed ownership information within natural texts, subsequently using Chain-of-Thought (CoT) prompting to construct semantically coherent and contextually natural QA pairs. This design ensures that fingerprints seamlessly integrate with the standard model behavior, remaining indistinguishable from regular outputs and substantially reducing the risk of accidental triggering and targeted removal. We conduct a comprehensive evaluation of ImF on 15 diverse LLMs, spanning different architectures and varying scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。