arXiv:2412.02323cs.CLcs.CR2024-12中稿 · ACL被引 6

针对藏文语言模型提出一种基于音节的文本对抗攻击方法

Pay Attention to the Robustness of Chinese Minority Language Models! Syllable-level Textual Adversarial Attack on Tibetan Script

  • 基于音节余弦距离与评分机制设计黑盒攻击
  • 在6个微调模型上成功生成高质量对抗样本
  • 揭示藏文NLP模型鲁棒性仍有巨大提升空间

文本对抗攻击通过精心设计的不可察觉扰动,使自然语言处理模型产生错误判断,也可用于评估模型鲁棒性。当前研究多集中于英语和中文,但对汉语少数民族语言的研究极少。本文针对藏文提出一种基于音节余弦距离与评分机制的黑盒文本对抗攻击方法TSAttacker。在两个预训练语言模型(PLM)微调的6个模型上,针对三个下游任务进行实验。结果表明,TSAttacker有效且能生成高质量对抗样本,同时暴露了现有模型鲁棒性仍有显著提升空间。

原文摘要 · Abstract (English)

The textual adversarial attack refers to an attack method in which the attacker adds imperceptible perturbations to the original texts by elaborate design so that the NLP (natural language processing) model produces false judgments. This method is also used to evaluate the robustness of NLP models. Currently, most of the research in this field focuses on English, and there is also a certain amount of research on Chinese. However, to the best of our knowledge, there is little research targeting Chinese minority languages. Textual adversarial attacks are a new challenge for the information processing of Chinese minority languages. In response to this situation, we propose a Tibetan syllable-level black-box textual adversarial attack called TSAttacker based on syllable cosine distance and scoring mechanism. And then, we conduct TSAttacker on six models generated by fine-tuning two PLMs (pre-trained language models) for three downstream tasks. The experiment results show that TSAttacker is effective and generates high-quality adversarial samples. In addition, the robustness of the involved models still has much room for improvement.

对抗攻击藏文NLP鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。