LoRA微调在对抗训练时攻击中存在安全短板,易受数据污染攻击。
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?

- 用神经正切核建模LoRA训练过程,分析其低秩结构的安全性
- 理论与实验证明LoRA对后门攻击更鲁棒,但对无目标数据投毒更脆弱
- 揭示低秩结构导致信息几何简化,是漏洞根源,适合模型安全研究者
低秩适配(LoRA)因其高效性成为大语言模型微调的主流方法。尽管已有大量研究关注其性能与结构特性,其在训练时攻击下的行为仍缺乏深入探讨,存在显著安全风险。本文从理论上研究了LoRA低秩结构在微调过程中对数据投毒和后门攻击的鲁棒性影响。我们提出一个分析框架,通过神经正切核简化训练动态分析,并结合信息论建立低秩结构与攻击脆弱性之间的联系。分析表明,相比全量微调,LoRA对后门攻击更具鲁棒性,但因信息几何过于简化,对无目标数据投毒更易受攻击。大量实验验证了理论发现。
原文摘要 · Abstract (English)
Low rank adaptation (LoRA) has emerged as a prominent technique for fine-tuning large language models (LLMs) thanks to its superb efficiency gains over previous methods. While extensive studies have examined the performance and structural properties of LoRA, its behavior upon training-time attacks remain underexplored, posing significant security risks. In this paper, we theoretically investigate the security implications of LoRA's low-rank structure during fine-tuning, in the context of its robustness against data poisoning and backdoor attacks. We propose an analytical framework that models LoRA's training dynamics, employs the neural tangent kernel to simplify the analysis of the training process, and applies information theory to establish connections between LoRA's low rank structure and its vulnerability against training-time attacks. Our analysis indicates that LoRA exhibits better robustness to backdoor attacks than full fine-tuning, while becomes more vulnerable to untargeted data poisoning due to its over-simplified information geometry. Extensive experimental evaluations have corroborated our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。