arXiv:2411.07843cs.CLcs.AI2024-11被引 1

利用汉字联想链生成对抗样本,让机器误解而人仍能懂。

Chain Association-based Attacking and Shielding Natural Language Processing Systems

  • 构建汉字联想图谱,生成可绕过AI的隐蔽干扰文本。
  • 在多个大模型上实现超90%攻击成功率,人类理解率仍超85%。
  • 提出训练与恢复双策略防御,适合安全敏感场景研究者。

联想使人类无需直白表达即可传递含义,但机器难以捕捉这种隐含关联。本文提出一种基于链式联想的自然语言处理系统对抗攻击方法,利用人机理解差异构建潜在对抗样本的搜索空间。首先基于汉字联想范式构建链式关联图谱,再采用离散粒子群优化算法搜索最优扰动文本。大量实验表明,包括大型语言模型在内的先进NLP系统均易受此攻击,而人类对扰动文本的理解准确率仍高于85%。同时探索了对抗训练和基于联想图的恢复两种防护策略。因涉及少量贬义词汇,部分内容可能引发不适。

原文摘要 · Abstract (English)

Association as a gift enables people do not have to mention something in completely straightforward words and allows others to understand what they intend to refer to. In this paper, we propose a chain association-based adversarial attack against natural language processing systems, utilizing the comprehension gap between humans and machines. We first generate a chain association graph for Chinese characters based on the association paradigm for building search space of potential adversarial examples. Then, we introduce an discrete particle swarm optimization algorithm to search for the optimal adversarial examples. We conduct comprehensive experiments and show that advanced natural language processing models and applications, including large language models, are vulnerable to our attack, while humans appear good at understanding the perturbed text. We also explore two methods, including adversarial training and associative graph-based recovery, to shield systems from chain association-based attack. Since a few examples that use some derogatory terms, this paper contains materials that may be offensive or upsetting to some people.

对抗攻击中文理解防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。