arXiv:2412.02343cs.CLcs.CR2024-12中稿 · WWW 2024 Workshop …被引 5

针对藏语文本提出多粒度对抗攻击方法,提升语言模型安全评估能力。

Multi-Granularity Tibetan Textual Adversarial Attack Method Based on Masked Language Model

  • 基于掩码语言模型生成候选替换词,按评分顺序实施攻击。
  • 使分类模型准确率下降超28.70%,90.60%样本预测被改变。
  • 首个面向藏语的文本对抗攻击方法,适合少数民族语言研究者。

在社交媒体中,神经网络模型广泛应用于仇恨言论检测、情感分析等任务,但其易受对抗攻击影响。例如,在文本分类任务中,攻击者通过微小扰动修改原文,几乎不改变语义却诱导模型产生错误预测。研究文本对抗攻击可评估并提升语言模型的鲁棒性。目前该领域主要聚焦英文和中文,对中文少数民族语言的研究较少。随着人工智能技术发展及少数民族语言模型的出现,文本对抗攻击成为中文少数民族语言信息处理的新挑战。为此,本文提出一种基于掩码语言模型的多粒度藏语文本对抗攻击方法TSTricker。利用掩码语言模型生成候选替换音节或词汇,采用评分机制确定替换顺序,并在多个微调的受害者模型上进行测试。实验结果表明,TSTricker使分类模型准确率降低超过28.70%,超过90.60%的样本预测被改变,攻击效果显著优于基线方法。

原文摘要 · Abstract (English)

In social media, neural network models have been applied to hate speech detection, sentiment analysis, etc., but neural network models are susceptible to adversarial attacks. For instance, in a text classification task, the attacker elaborately introduces perturbations to the original texts that hardly alter the original semantics in order to trick the model into making different predictions. By studying textual adversarial attack methods, the robustness of language models can be evaluated and then improved. Currently, most of the research in this field focuses on English, and there is also a certain amount of research on Chinese. However, there is little research targeting Chinese minority languages. With the rapid development of artificial intelligence technology and the emergence of Chinese minority language models, textual adversarial attacks become a new challenge for the information processing of Chinese minority languages. In response to this situation, we propose a multi-granularity Tibetan textual adversarial attack method based on masked language models called TSTricker. We utilize the masked language models to generate candidate substitution syllables or words, adopt the scoring mechanism to determine the substitution order, and then conduct the attack method on several fine-tuned victim models. The experimental results show that TSTricker reduces the accuracy of the classification models by more than 28.70% and makes the classification models change the predictions of more than 90.60% of the samples, which has an evidently higher attack effect than the baseline method.

藏语处理对抗攻击多粒度语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。