十种主流模型指纹方案均易被恶意攻击绕过,需设计更鲁棒的防御机制。
Are Robust LLM Fingerprints Adversarially Robust?
- 针对指纹方案漏洞设计针对性对抗攻击
- 十种指纹方案均被完全绕过且模型性能不受影响
- 提醒指纹设计者必须内置对抗鲁棒性
模型指纹技术被视为声明模型所有权的有力手段。然而,现有评估多聚焦于良性扰动(如增量微调、模型合并、提示注入),缺乏对恶意模型托管方发起的对抗鲁棒性系统的系统研究。为此,我们首先定义了一个具体且实用的威胁模型。接着,深入分析了现有指纹方案的根本弱点,并据此开发出适配各漏洞的自适应对抗攻击。实验表明,这些攻击能完全绕过近期提出的十种指纹方案,同时保持模型对终端用户的高可用性。本工作呼吁指纹设计者将对抗鲁棒性作为核心考量,最后提出未来指纹方法的设计建议。
原文摘要 · Abstract (English)
Model fingerprinting has emerged as a promising paradigm for claiming model ownership. However, robustness evaluations of these schemes have mostly focused on benign perturbations such as incremental fine-tuning, model merging, and prompting. Lack of systematic investigations into {\em adversarial robustness} against a malicious model host leaves current systems vulnerable. To bridge this gap, we first define a concrete, practical threat model against model fingerprinting. We then take a critical look at existing model fingerprinting schemes to identify their fundamental vulnerabilities. Based on these, we develop adaptive adversarial attacks tailored for each vulnerability, and demonstrate that these can bypass model authentication completely for ten recently proposed fingerprinting schemes while maintaining high utility of the model for the end users. Our work encourages fingerprint designers to adopt adversarial robustness by design. We end with recommendations for future fingerprinting methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。