arXiv:2602.09434cs.CRcs.AI2026-02被引 2

用拒绝向量追踪大模型来源,防抄袭更可靠

A Behavioral Fingerprint for Large Language Models: Provenance Tracking via Refusal Vectors

  • 从模型对有害/无害提示的响应差异中提取拒绝向量作为行为指纹
  • 在76个衍生模型中100%准确识别母本模型家族,抗微调和量化攻击
  • 支持隐私保护下的公开验证,适合模型版权保护场景

由于未经授权的衍生模型泛滥,保护大语言模型(LLMs)知识产权成为关键挑战。本文提出一种新型指纹框架,利用安全对齐引发的行为模式,通过拒绝向量实现LLM溯源。这些向量源自模型在处理有害与无害提示时内部表示的方向性差异,构成鲁棒的行为指纹。我们构建了基于该概念的指纹系统,并进行了广泛验证,证明其对微调、合并和量化等常见修改具有高度鲁棒性。实验显示,不同独立训练的模型间指纹余弦相似度低,每组指纹唯一。在涵盖76个后代模型的大规模识别任务中,方法达到100%准确率。此外,分析表明,在对齐破坏攻击下性能虽下降,但可检测痕迹仍存。最后,我们提出理论框架,通过局部敏感哈希与零知识证明,将私有指纹转化为可公开验证且保护隐私的凭证。

原文摘要 · Abstract (English)

Protecting the intellectual property of large language models (LLMs) is a critical challenge due to the proliferation of unauthorized derivative models. We introduce a novel fingerprinting framework that leverages the behavioral patterns induced by safety alignment, applying the concept of refusal vectors for LLM provenance tracking. These vectors, extracted from directional patterns in a model's internal representations when processing harmful versus harmless prompts, serve as robust behavioral fingerprints. Our contribution lies in developing a fingerprinting system around this concept and conducting extensive validation of its effectiveness for IP protection. We demonstrate that these behavioral fingerprints are highly robust against common modifications, including finetunes, merges, and quantization. Our experiments show that the fingerprint is unique to each model family, with low cosine similarity between independently trained models. In a large-scale identification task across 76 offspring models, our method achieves 100\% accuracy in identifying the correct base model family. Furthermore, we analyze the fingerprint's behavior under alignment-breaking attacks, finding that while performance degrades significantly, detectable traces remain. Finally, we propose a theoretical framework to transform this private fingerprint into a publicly verifiable, privacy-preserving artifact using locality-sensitive hashing and zero-knowledge proofs.

模型溯源行为指纹知识产权零知识证明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。