用交叉熵优化提升文本对抗攻击的通用性与隐蔽性
Cross-Entropy Attacks to Language Models via Rare Event Simulation
- 基于交叉熵优化,统一处理软标签与硬标签攻击目标
- 在文本分类和翻译任务中实现更高攻击成功率与更优语句质量
- 适合研究模型鲁棒性或对抗样本生成的从业者使用
黑箱文本对抗攻击因缺乏模型信息及文本离散不可导特性而困难重重。现有方法往往泛化能力弱,受限于词重要性排序的低效优化,且常以牺牲语义完整性换取攻击效果。本文提出一种新方法——交叉熵攻击(CEA),利用交叉熵优化解决上述问题。该方法在软标签与硬标签场景下均定义了对抗目标,并通过交叉熵优化识别最优替换项。在文档分类与语言翻译任务上的大量实验表明,该攻击方法在攻击性能、不可察觉性和句子质量方面均表现优异。
原文摘要 · Abstract (English)
Black-box textual adversarial attacks are challenging due to the lack of model information and the discrete, non-differentiable nature of text. Existing methods often lack versatility for attacking different models, suffer from limited attacking performance due to the inefficient optimization with word saliency ranking, and frequently sacrifice semantic integrity to achieve better attack outcomes. This paper introduces a novel approach to textual adversarial attacks, which we call Cross-Entropy Attacks (CEA), that uses Cross-Entropy optimization to address the above issues. Our CEA approach defines adversarial objectives for both soft-label and hard-label settings and employs CE optimization to identify optimal replacements. Through extensive experiments on document classification and language translation problems, we demonstrate that our attack method excels in terms of attacking performance, imperceptibility, and sentence quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。