arXiv:2502.18511cs.CRcs.AI2025-02ACL被引 14

构建首个高效后门攻击评测基准,覆盖12种攻击方法与12个大模型。

ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models

  • 支持参数高效微调与无微调注入后门,评估更贴近真实场景。
  • 实测显示参数高效攻击在分类任务中更优,且跨数据集泛化能力强。
  • 提供标准化工具箱,适合安全研究者与模型防护开发者使用。

生成式大语言模型在自然语言处理中至关重要,但易受后门攻击影响,即通过细微触发器改变其行为。尽管针对大模型的后门攻击不断涌现,现有评测基准在攻击覆盖面、指标完整性及攻击对齐性方面仍显不足,且预训练攻击在实际中因资源限制而理想化。为此,我们提出ELBA-Bench,一个全面统一的框架,支持通过参数高效微调(如LoRA)或无需微调技术(如上下文学习)注入后门。ELBA-Bench涵盖超过1300次实验,实现12种攻击方法、18个数据集和12个大模型的系统评估。大量实验揭示了各类攻击策略的优劣:参数高效微调攻击在分类任务中持续优于无微调方法,并在优化触发器后展现出强跨数据集泛化能力;任务相关的后门优化技术或攻击提示结合干净与对抗样本演示,可提升攻击成功率,同时保持对干净样本的性能。此外,我们还引入一个通用工具箱,旨在推动该关键领域的标准化研究进展。

原文摘要 · Abstract (English)

Generative large language models are crucial in natural language processing, but they are vulnerable to backdoor attacks, where subtle triggers compromise their behavior. Although backdoor attacks against LLMs are constantly emerging, existing benchmarks remain limited in terms of sufficient coverage of attack, metric system integrity, backdoor attack alignment. And existing pre-trained backdoor attacks are idealized in practice due to resource access constraints. Therefore we establish $\textit{ELBA-Bench}$, a comprehensive and unified framework that allows attackers to inject backdoor through parameter efficient fine-tuning ($\textit{e.g.,}$ LoRA) or without fine-tuning techniques ($\textit{e.g.,}$ In-context-learning). $\textit{ELBA-Bench}$ provides over 1300 experiments encompassing the implementations of 12 attack methods, 18 datasets, and 12 LLMs. Extensive experiments provide new invaluable findings into the strengths and limitations of various attack strategies. For instance, PEFT attack consistently outperform without fine-tuning approaches in classification tasks while showing strong cross-dataset generalization with optimized triggers boosting robustness; Task-relevant backdoor optimization techniques or attack prompts along with clean and adversarial demonstrations can enhance backdoor attack success while preserving model performance on clean samples. Additionally, we introduce a universal toolbox designed for standardized backdoor attack research, with the goal of propelling further progress in this vital area.

后门攻击大模型安全评测基准参数高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。