合成数据可传播恶意攻击,新框架VIA让毒化内容在生成中扩散
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
- 将毒化载荷藏于保护壳中,借良性样本劫持点生成恶意合成数据
- 在纯干净查询下,攻击成功率提升至与上游污染模型相当水平
- 揭示合成数据训练范式中的安全漏洞,适合关注LLM安全的研究者
合成数据指由模型生成的伪样本。尽管其已被证实能显著提升大语言模型(LLMs)训练性能并广泛应用于开发,但可能引入的安全风险尚未被系统研究。本文系统评估了集成合成数据训练范式对主流投毒和后门攻击的鲁棒性。结果表明,该范式对现有攻击具有较强抗性,主要源于毒化数据与生成合成样本所用查询之间的分布差异。为提升攻击有效性并进一步探究合成数据带来的安全风险,我们提出一种新颖且通用的攻击框架——病毒感染攻击(VIA),可在仅使用纯净查询的情况下,实现攻击通过合成数据传播。受网络安全中病毒设计原理启发,VIA将毒化载荷隐藏于保护壳中,并在良性样本中寻找最优劫持点以最大化生成恶意内容的概率。在数据投毒和后门攻击上的大量实验表明,VIA显著增加了合成数据中毒化内容的比例,并使下游模型的攻击成功率达到与上游污染模型相当的水平。
原文摘要 · Abstract (English)
Synthetic data refers to artificial samples generated by models. While it has been validated to significantly enhance the performance of large language models (LLMs) during training and has been widely adopted in LLM development, potential security risks it may introduce remain uninvestigated. This paper systematically evaluates the resilience of synthetic-data-integrated training paradigm for LLMs against mainstream poisoning and backdoor attacks. We reveal that such a paradigm exhibits strong resistance to existing attacks, primarily thanks to the different distribution patterns between poisoning data and queries used to generate synthetic samples. To enhance the effectiveness of these attacks and further investigate the security risks introduced by synthetic data, we introduce a novel and universal attack framework, namely, Virus Infection Attack (VIA), which enables the propagation of current attacks through synthetic data even under purely clean queries. Inspired by the principles of virus design in cybersecurity, VIA conceals the poisoning payload within a protective "shell" and strategically searches for optimal hijacking points in benign samples to maximize the likelihood of generating malicious content. Extensive experiments on both data poisoning and backdoor attacks show that VIA significantly increases the presence of poisoning content in synthetic data and correspondingly raises the attack success rate (ASR) on downstream models to levels comparable to those observed in the poisoned upstream models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。