利用标题偏倚和影响函数攻击摘要模型,使其生成提取式摘要。
Attacks against Abstractive Text Summarization Models through Lead Bias and Influence Functions
- 通过挖掘摘要模型的标题偏好,实施对抗性扰动。
- 引入影响函数实现数据投毒,破坏模型完整性。
- 攻击使模型从抽象摘要转为提取式摘要,适合安全研究者参考。
大型语言模型在文本理解和生成方面带来了新机遇,但其在文本分类和翻译任务中易受对抗性扰动和数据投毒攻击的影响。然而,抽象式文本摘要模型的对抗鲁棒性尚未被充分探索。本文揭示了一种新方法:利用摘要模型固有的标题偏倚进行对抗性扰动,并创新性地应用影响函数执行数据投毒,损害模型完整性。该方法不仅导致模型行为出现偏差,使其产生预期输出,还引发新行为变化——被攻击的模型倾向于生成提取式摘要而非抽象式摘要。
原文摘要 · Abstract (English)
Large Language Models have introduced novel opportunities for text comprehension and generation. Yet, they are vulnerable to adversarial perturbations and data poisoning attacks, particularly in tasks like text classification and translation. However, the adversarial robustness of abstractive text summarization models remains less explored. In this work, we unveil a novel approach by exploiting the inherent lead bias in summarization models, to perform adversarial perturbations. Furthermore, we introduce an innovative application of influence functions, to execute data poisoning, which compromises the model's integrity. This approach not only shows a skew in the models behavior to produce desired outcomes but also shows a new behavioral change, where models under attack tend to generate extractive summaries rather than abstractive summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。