arXiv:2502.14828cs.LGcs.CR2025-02NeurIPS被引 13

攻击者利用模型输出的自然变化隐蔽传递有害知识,绕过逐样本检测。

Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs

  • 用良性输出中的语义/语法变异伪装有害信息,实现逐样本不可检测。
  • 在OpenAI API上成功诱导模型回答有害多选题,且避开强化监控系统。
  • 适合关注大模型安全与细调防御漏洞的研究者阅读。

大模型开发者通过技术手段防范细调滥用攻击——即攻击者通过公开API微调模型以规避安全机制。此前研究已针对特定防御提出有效攻击。本文揭示:仅依赖逐样本检测有害训练或推理样本的防御存在根本局限。我们构造出‘逐样本不可检测’的攻击,利用良性模型输出中的熵(如语义或语法变化)隐蔽传递危险知识。攻击所用样本均为事先从模型获取的无害良性数据,训练与推理样本均看似正常且低困惑度。在OpenAI细调API上的测试表明,该攻击可成功诱导模型回答有害多选题,且能逃避我们设计的、能检测其他细调攻击的增强监控系统。我们呼吁社区发展应对此类根本性缺陷的新防御策略。

原文摘要 · Abstract (English)

LLM developers have imposed technical interventions to prevent fine-tuning misuse attacks, attacks where adversaries evade safeguards by fine-tuning the model using a public API. Previous work has established several successful attacks against specific fine-tuning API defences. In this work, we show that defences of fine-tuning APIs that seek to detect individual harmful training or inference samples ('pointwise' detection) are fundamentally limited in their ability to prevent fine-tuning attacks. We construct 'pointwise-undetectable' attacks that repurpose entropy in benign model outputs (e.g. semantic or syntactic variations) to covertly transmit dangerous knowledge. Our attacks are composed solely of unsuspicious benign samples that can be collected from the model before fine-tuning, meaning training and inference samples are all individually benign and low-perplexity. We test our attacks against the OpenAI fine-tuning API, finding they succeed in eliciting answers to harmful multiple-choice questions, and that they evade an enhanced monitoring system we design that successfully detects other fine-tuning attacks. We encourage the community to develop defences that tackle the fundamental limitations we uncover in pointwise fine-tuning API defences.

大模型安全细调攻击对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。