发现大模型参数分布指纹,可追踪模型来源并识别抄袭。
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
- 利用注意力层参数标准差分布作为稳定指纹。
- 即使持续训练后仍能准确识别模型血缘关系。
- 适合模型版权保护与抄袭检测场景。
大语言模型面临日益严峻的版权与知识产权挑战,随着训练成本上升和模型复用普遍化,现有水印技术对持续训练不鲁棒,难以保障模型归属。本文提出一种基于模型内在特性的鲁棒指纹方法:发现不同层注意力参数矩阵的标准差分布具有独特且稳定的模式,即便经过大量持续训练仍保持不变。这些参数分布特征可作为可靠指纹,用于模型身份认证与版权侵权检测。实验验证了该方法在多个模型家族中的有效性。特别地,研究发现华为最新发布的Pangu Pro MoE模型极可能通过升级技术源自Qwen-2.5 14B模型,而非从头训练,揭示潜在模型剽窃、版权侵犯与信息伪造问题。结果表明,仅靠刻意持续训练无法彻底隐藏模型来源,亟需发展更可靠的指纹保护机制。
原文摘要 · Abstract (English)
Large language models (LLMs) face significant copyright and intellectual property challenges as the cost of training increases and model reuse becomes prevalent. While watermarking techniques have been proposed to protect model ownership, they may not be robust to continue training and development, posing serious threats to model attribution and copyright protection. This work introduces a simple yet effective approach for robust LLM fingerprinting based on intrinsic model characteristics. We discover that the standard deviation distributions of attention parameter matrices across different layers exhibit distinctive patterns that remain stable even after extensive continued training. These parameter distribution signatures serve as robust fingerprints that can reliably identify model lineage and detect potential copyright infringement. Our experimental validation across multiple model families demonstrates the effectiveness of our method for model authentication. Notably, our investigation uncovers evidence that a recently Pangu Pro MoE model released by Huawei is derived from Qwen-2.5 14B model through upcycling techniques rather than training from scratch, highlighting potential cases of model plagiarism, copyright violation, and information fabrication. These findings underscore the critical importance of developing robust fingerprinting methods for protecting intellectual property in large-scale model development and emphasize that deliberate continued training alone is insufficient to completely obscure model origins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。