黑客可用低成本伪造强模型指纹,骗过用户验证。
Your "Pro" LLM Subscription May Actually Be "Free": Exposing Fingerprint Spoofing Risks in LLM Inference Services

- 用微调弱模型模仿强模型特征,骗过指纹检测
- 实测可绕过主流指纹方法,且微调成本极低
- 适合关注AI服务可信性的研究人员和开发者
随着大语言模型(LLM)API日益普及,用户依赖黑箱指纹技术验证服务是否提供宣称的高级模型。然而,当前方法可能忽略恶意提供商通过参数高效微调弱模型以模仿强模型的行为。本文提出新型威胁‘指纹欺骗’:攻击者在不被察觉的情况下,使用弱模型伪装成强模型,从而规避用户端指纹验证。我们首先从理论上证明,因用户资源有限(查询预算有限、指纹分类器较弱),现有指纹方法易受此类攻击。基于此分析,我们提出GhostPrint攻击框架,结合代理建模、奖励排序微调与知识蒸馏,实现低成本欺骗。在静态与持续指纹场景下的大量实验表明,该方法能使弱模型持续绕过主流指纹检测,同时保持良好实用性,暴露当前LLM指纹验证流程中的关键漏洞。
原文摘要 · Abstract (English)
As Large Language Model (LLM) APIs become ubiquitous, users increasingly rely on black-box fingerprinting to verify that providers are serving the advertised premium models. However, these methods may overlook adversarial providers who manipulate model weights to cheat the fingerprint process. We introduce a novel threat termed fingerprint spoofing, where a malicious provider stealthily serves a weaker model that has been parameter-efficiently fine-tuned to mimic a stronger model, thereby evading user-side fingerprinting. We first formally prove that user-side resource constraints (i.e., finite query budgets and weak fingerprinting classifiers) make current fingerprinting vulnerable to fingerprint spoofing. Guided by this theoretical analysis, we propose GhostPrint, a cost-effective attack framework leveraging surrogate modeling, reward-ranked fine-tuning, and knowledge distillation. Extensive evaluations in both static and continual fingerprinting settings demonstrate that GhostPrint allows weak models to consistently bypass representative fingerprint methods while maintaining utility at a low fine-tuning cost, exposing a critical vulnerability in current LLM fingerprinting pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。