提出一种高效擦除大模型指纹的方法,保护模型不被非法追踪。
MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
- 用错配与干净数据分两阶段微调,精准去除指纹。
- 仅需少于1000样本即可完全移除指纹且保持性能。
- 支持跨模型迁移,无需重复训练,适合安全研究者使用。
大语言模型在多个领域广泛应用,引发模型所有权和知识产权保护的关切。尽管基于后门的指纹技术成为模型认证的有前景方案,但有效擦除这些指纹的方法仍基本未被探索。为此,我们提出一种名为不匹配擦除器(MEraser)的新方法,可在保持模型性能的前提下,有效去除大语言模型中的后门指纹。该方法采用两阶段微调策略,利用精心构建的错配数据集与干净数据集。在多个大语言模型架构及指纹方法上的广泛评估表明,MEraser仅需少于1,000个样本即可实现指纹的完全移除,并保持模型性能。此外,我们还提出可迁移的擦除机制,实现跨模型的指纹移除而无需重复训练。本工作不仅为大模型指纹擦除提供了实用方案,揭示了当前指纹技术的关键漏洞,还建立了全面的评估基准,为未来开发更鲁棒的模型保护方法奠定基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have become increasingly prevalent across various sectors, raising critical concerns about model ownership and intellectual property protection. Although backdoor-based fingerprinting has emerged as a promising solution for model authentication, effective attacks for removing these fingerprints remain largely unexplored. Therefore, we present Mismatched Eraser (MEraser), a novel method for effectively removing backdoor-based fingerprints from LLMs while maintaining model performance. Our approach leverages a two-phase fine-tuning strategy utilizing carefully constructed mismatched and clean datasets. Through extensive evaluation across multiple LLM architectures and fingerprinting methods, we demonstrate that MEraser achieves complete fingerprinting removal while maintaining model performance with minimal training data of fewer than 1,000 samples. Furthermore, we introduce a transferable erasure mechanism that enables effective fingerprinting removal across different models without repeated training. In conclusion, our approach provides a practical solution for fingerprinting removal in LLMs, reveals critical vulnerabilities in current fingerprinting techniques, and establishes comprehensive evaluation benchmarks for developing more resilient model protection methods in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。