arXiv:2508.09021cs.CRcs.AI2025-08被引 1

用强化学习提升大模型指纹攻击效率,同时提出保语义的防御方案。

Attacks and Defenses Against LLM Fingerprinting

  • 用强化学习自动选查询,3次即可精准识别模型
  • 防御方用第二模型过滤输出,使指纹识别率下降且保持语义质量
  • 兼顾攻击与防御,适合关注模型隐私安全的研究者

随着大语言模型在敏感环境中的部署增多,指纹攻击带来显著的隐私与安全风险。本文从攻防双重视角开展研究:攻击方面,利用强化学习自动优化查询选择,在仅使用3个查询的情况下,相比随机选取相同数量查询,显著提升了指纹识别准确率;防御方面,通过一个二级大语言模型进行语义保持的输出过滤,有效掩盖模型身份,降低指纹识别成功率,同时维持输出质量。实验表明该方法在多个测试模型上均能有效抑制指纹攻击。本研究揭示了指纹技术的潜力,同时也提供了切实可行的对抗策略。

原文摘要 · Abstract (English)

As large language models are increasingly deployed in sensitive environments, fingerprinting attacks pose significant privacy and security risks. We present a study of LLM fingerprinting from both offensive and defensive perspectives. Our attack methodology uses reinforcement learning to automatically optimize query selection, achieving better fingerprinting accuracy with only 3 queries compared to randomly selecting 3 queries from the same pool. Our defensive approach employs semantic-preserving output filtering through a secondary LLM to obfuscate model identity while maintaining semantic integrity. The defensive method reduces fingerprinting accuracy across tested models while preserving output quality. These contributions show the potential to improve fingerprinting tools capabilities while providing practical mitigation strategies against fingerprinting attacks.

模型指纹隐私安全强化学习防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。