arXiv:2501.00066cs.CLcs.AI2025-01

大模型在迁移学习中易受攻击,模型越大越抗干扰。

On Adversarial Robustness of Language Models in Transfer Learning

  • 对比多种模型在迁移学习中的抗干扰能力。
  • 大模型比小模型更耐攻击,但迁移学习普遍降低安全性。
  • 适合关注大模型安全性的研究者和应用开发者。

我们研究了大型语言模型在迁移学习场景下的对抗鲁棒性。通过对多个数据集(MBIB Hate Speech、MBIB Political Bias、MBIB Gender Bias)和多种模型架构(BERT、RoBERTa、GPT-2、Gemma、Phi)的全面实验发现,尽管迁移学习能提升标准性能指标,却常导致模型对对抗攻击更加脆弱。研究结果表明,更大规模的模型表现出更强的鲁棒性,揭示了模型大小、架构与适配方法之间复杂的相互作用。本工作强调在迁移学习中必须考虑对抗鲁棒性,为在不牺牲性能的前提下保障模型安全提供了重要洞见。这些发现对实际应用中同时追求高性能与高安全性的大模型开发具有重要意义。

原文摘要 · Abstract (English)

We investigate the adversarial robustness of LLMs in transfer learning scenarios. Through comprehensive experiments on multiple datasets (MBIB Hate Speech, MBIB Political Bias, MBIB Gender Bias) and various model architectures (BERT, RoBERTa, GPT-2, Gemma, Phi), we reveal that transfer learning, while improving standard performance metrics, often leads to increased vulnerability to adversarial attacks. Our findings demonstrate that larger models exhibit greater resilience to this phenomenon, suggesting a complex interplay between model size, architecture, and adaptation methods. Our work highlights the crucial need for considering adversarial robustness in transfer learning scenarios and provides insights into maintaining model security without compromising performance. These findings have significant implications for the development and deployment of LLMs in real-world applications where both performance and robustness are paramount.

大模型对抗攻击迁移学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。