揭露大模型微调中的安全漏洞及防护策略
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
- 系统梳理有害微调攻击的威胁模型与变体
- 总结主流攻击手段与防御方法的机制差异
- 适合关注AI安全与可信部署的研究者参考
近期研究揭示,新兴的微调即服务商业模式存在严重安全隐患:用户上传少量有害数据即可破坏模型的安全对齐。这种称为有害微调攻击的行为已在学术界和工业界引发广泛关注。本文首先系统化地定义了该问题的威胁模型与基本假设,随后从三个核心维度全面综述:攻击设置、防御设计与评估方法。首先阐述问题的威胁模型,介绍有害微调攻击及其多种变体;其次系统梳理现有文献中代表性攻击、防御方法及副作用机理分析;最后介绍评估方法并展望未来研究方向,为该领域未来发展提供指南与关键视角。我们还维护一份精选论文列表,可通过 https://github.com/git-disl/awesome_LLM-harmful-fine-tuning-papers 获取。
原文摘要 · Abstract (English)
Recent research demonstrates that the nascent fine-tuning-as-a-service business model exposes serious safety concerns: fine-tuning with a few harmful data uploaded from the users can compromise the safety alignment of the model. The attack, known as harmful fine-tuning attack, has generated broad research interests in both academia and industry. In this paper, we first systematically formulate the threat model and basic assumptions of harmful fine-tuning. Then, we provide a comprehensive review of harmful fine-tuning from three fundamental perspectives: attack setting, defense design, and evaluation methodology. First, we present the threat model of the problem and introduce the harmful fine-tuning attack and its variants. Next, we systematically survey representative attacks, defense methods, and mechanical analysis of adverse effects in the existing literature. Finally, we introduce the evaluation methodology and outline future research directions, which can serve as guidelines and crucial perspectives for the future development of the subject. We also maintain a curated list of relevant papers, which are made accessible at https://github.com/git-disl/awesome_LLM-harmful-fine-tuning-papers
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。