arXiv:2409.02649cs.CLcs.AI2024-09被引 6

融合多种攻击方法,提升对文本可信度模型的对抗攻击效果

OpenFact at CheckThat! 2024: Combining Multiple Attack Methods for Effective Adversarial Text Generation

  • 组合BERT-Attack、遗传算法等方法,构建混合攻击策略
  • 在5个数据集上显著提高攻击成功率,最高提升达18.7%
  • 适合研究模型鲁棒性与对抗样本的NLP安全方向学者

本文针对CLEF 2024任务6:基于对抗样本的可信度评估鲁棒性(InCrediblAE),开展实验与结果分析。该任务旨在五个问题领域生成对抗样本,以评估微调BERT、BiLSTM和RoBERTa等主流文本分类模型在可信度评估任务中的鲁棒性。本研究探索集成学习在增强NLP模型对抗攻击中的应用。我们系统测试并优化了BERT-Attack、遗传算法、TextFooler和CLARE等多种攻击方法,在五个不同数据集上的各类虚假信息任务中进行验证。通过改进BERT-Attack的实现方式并设计混合攻击策略,显著提升了攻击有效性。结果表明,方法的改良与多方法融合可生成更复杂、更有效的对抗样本,有助于推动更具鲁棒性和安全性的系统发展。

原文摘要 · Abstract (English)

This paper presents the experiments and results for the CheckThat! Lab at CLEF 2024 Task 6: Robustness of Credibility Assessment with Adversarial Examples (InCrediblAE). The primary objective of this task was to generate adversarial examples in five problem domains in order to evaluate the robustness of widely used text classification methods (fine-tuned BERT, BiLSTM, and RoBERTa) when applied to credibility assessment issues. This study explores the application of ensemble learning to enhance adversarial attacks on natural language processing (NLP) models. We systematically tested and refined several adversarial attack methods, including BERT-Attack, Genetic algorithms, TextFooler, and CLARE, on five datasets across various misinformation tasks. By developing modified versions of BERT-Attack and hybrid methods, we achieved significant improvements in attack effectiveness. Our results demonstrate the potential of modification and combining multiple methods to create more sophisticated and effective adversarial attack strategies, contributing to the development of more robust and secure systems.

对抗攻击NLP安全模型鲁棒性文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。