arXiv:2505.12682cs.LG2025-05被引 9

通过罕见提示区域稳定特征,实现大模型血统识别。

RAFP: Identifying LLM Lineages via Rare-Region Fingerprints

  • 利用低概率提示区域的梯度优化构建非侵入式指纹。
  • 在微调、量化等操作后仍保持指纹有效性,准确率超基线。
  • 适合模型版权验证,尤其适用于黑盒场景。

大型语言模型(LLMs)正越来越多地以受限许可证发布,亟需可靠的模型所有权验证方法。现有指纹技术在下游微调下易失效,需侵入式训练修改,或在黑盒设置中表现不佳。我们提出RAFP,一种基于罕见区域指纹的鲁棒模型血统识别框架。核心思想是:下游微调主要更新高频语言行为,而低概率提示区域受优化信号弱、梯度对齐有限,其响应行为在常见模型适配中保持稳定。RAFP无需修改模型权重,通过离散梯度优化罕见提示生成指纹。理论分析表明,微调下罕见区域指纹的似然变化有界。在四个LLM家族及多种下游适应(监督微调、LoRA、量化、提示模板变化、解码策略调整)上的实验显示,RAFP在黑盒场景中显著优于现有基线,指纹持久性强。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly released under restricted licenses, creating a growing need for robust model ownership verification. Existing fingerprinting methods are often fragile under downstream finetuning, require invasive training modifications, or fail in black-box settings. We introduce RAFP, a robust framework for identifying LLM lineages via rare-region fingerprints. Our key insight is that downstream finetuning primarily updates common high-density language behaviors, while low-probability prompt regions receive weak optimization signal and limited gradient alignment under finetuned distribution. As a result, rare prompt-response behaviors remain stable across common model adaptations. RAFP is non-invasive, constructing fingerprints via discrete gradient-based optimization over rare prompts without modifying model weights. We provide a theoretical analysis showing that the likelihood change of rare-region fingerprints under finetuning remains bounded. Experiments across four LLM families and multiple downstream adaptations, including supervised finetuning, LoRA, quantization, prompt-template variation, and decoding changes, show that RAFP achieves strong fingerprint persistence and substantially outperforms prior fingerprinting baselines in black-box settings.

模型指纹大模型版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。