用大模型生成更自然的对抗样本,提升NLP安全测试效果。
Advancing NLP Security by Leveraging LLMs as Adversarial Engines
- 用大模型自动生成多样化的对抗攻击,包括扰动和恶意补丁。
- 生成的攻击在语义上更连贯,能绕过多种分类器防御。
- 适合安全研究者、模型开发者关注新型漏洞与防御策略。
本文提出一种新方法,利用大语言模型(LLMs)作为生成对抗攻击的引擎,推动NLP安全研究。基于近期研究表明大模型可有效生成词级对抗样本,本文主张将其扩展至更广泛的攻击类型,包括对抗补丁、通用扰动和目标攻击。凭借其强大的语言理解与生成能力,大模型能生成更具有效性、语义连贯性及人类风格的对抗样本,适用于多种领域与模型架构。这一范式转变具有深远意义,有望增强模型鲁棒性、揭示新漏洞,并推动防御技术革新。通过探索这一前沿方向,旨在促进关键应用中更安全、可靠、可信的NLP系统发展。
原文摘要 · Abstract (English)
This position paper proposes a novel approach to advancing NLP security by leveraging Large Language Models (LLMs) as engines for generating diverse adversarial attacks. Building upon recent work demonstrating LLMs' effectiveness in creating word-level adversarial examples, we argue for expanding this concept to encompass a broader range of attack types, including adversarial patches, universal perturbations, and targeted attacks. We posit that LLMs' sophisticated language understanding and generation capabilities can produce more effective, semantically coherent, and human-like adversarial examples across various domains and classifier architectures. This paradigm shift in adversarial NLP has far-reaching implications, potentially enhancing model robustness, uncovering new vulnerabilities, and driving innovation in defense mechanisms. By exploring this new frontier, we aim to contribute to the development of more secure, reliable, and trustworthy NLP systems for critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。