arXiv:2505.16530cs.CRcs.AI2025-05Conference of the …被引 11

提出双层指纹框架,实现黑盒下大模型版权精准验证

DuFFin: A Dual-Level Fingerprinting Framework for LLMs IP Protection

  • 通过触发模式与知识级指纹双重特征识别模型来源
  • 在多种微调/量化版本上实现0.95以上版权验证准确率
  • 适用于开源模型版权保护,适合企业与开发者使用

大语言模型因其高昂的训练成本被视为重要知识产权。为防止恶意盗用或未经授权部署,亟需有效的保护机制。现有水印与指纹技术或影响生成质量,或受限于白盒访问,实用性不足。为此,我们提出DuFFin——一种面向黑盒环境的双层指纹框架,通过提取触发模式和知识级指纹来识别可疑模型来源。我们在多个开源平台收集的模型上进行实验,涵盖四种主流基础模型及其微调、量化、安全对齐等变体,这些模型由大型公司、初创企业和个人发布。结果表明,该方法可准确验证基础模型的版权归属,在其各类衍生版本上均达到超过0.95的IP-ROC指标。代码已公开于https://github.com/yuliangyan0807/llm-fingerprint。

原文摘要 · Abstract (English)

Large language models (LLMs) are considered valuable Intellectual Properties (IP) for legitimate owners due to the enormous computational cost of training. It is crucial to protect the IP of LLMs from malicious stealing or unauthorized deployment. Despite existing efforts in watermarking and fingerprinting LLMs, these methods either impact the text generation process or are limited in white-box access to the suspect model, making them impractical. Hence, we propose DuFFin, a novel $\textbf{Du}$al-Level $\textbf{Fin}$gerprinting $\textbf{F}$ramework for black-box setting ownership verification. DuFFin extracts the trigger pattern and the knowledge-level fingerprints to identify the source of a suspect model. We conduct experiments on a variety of models collected from the open-source website, including four popular base models as protected LLMs and their fine-tuning, quantization, and safety alignment versions, which are released by large companies, start-ups, and individual users. Results show that our method can accurately verify the copyright of the base protected LLM on their model variants, achieving the IP-ROC metric greater than 0.95. Our code is available at https://github.com/yuliangyan0807/llm-fingerprint.

模型指纹版权保护大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。