arXiv:2410.04377cs.LGcs.CL2024-10中稿 · in Computational L…

提出量化文本对抗样本可疑度的新方法,提升欺骗性。

Graded Suspiciousness of Adversarial Texts to Human

  • 构建人类对对抗文本可疑度的评分数据集,基于四种攻击方法
  • 建立回归模型预测可疑度,相关性达0.78(皮尔逊)
  • 可嵌入生成流程,降低文本被识破概率,适合安全与内容生成研究者

对抗样本对图像和文本领域的深度神经网络构成重大挑战,其通过精心修改输入来降低模型性能。与图像对抗样本不同,文本对抗样本需保持语义相似性且具有离散性。本研究聚焦人类对对抗文本的‘可疑感’,不同于图像中追求不可察觉的隐秘性,文本对抗内容需在欺骗自然语言系统的同时避免引起人类读者怀疑。我们构建并公开了一个新型的利克特量表人类评估数据集,涵盖四种主流攻击方法生成的对抗句,并分析其与人类检测机器修改能力的相关性。进一步,我们开发了一种基于回归的可疑度量化模型,建立了未来研究基准。实验表明,该模型能有效预测人类判断,相关系数达0.78。此外,我们将可疑度分数融入生成过程,使生成文本更难被识别为机器生成。相关数据与代码已公开。

原文摘要 · Abstract (English)

Adversarial examples pose a significant challenge to deep neural networks (DNNs) across both image and text domains, with the intent to degrade model performance through meticulously altered inputs. Adversarial texts, however, are distinct from adversarial images due to their requirement for semantic similarity and the discrete nature of the textual contents. This study delves into the concept of human suspiciousness, a quality distinct from the traditional focus on imperceptibility found in image-based adversarial examples. Unlike images, where adversarial changes are meant to be indistinguishable to the human eye, textual adversarial content must often remain undetected or non-suspicious to human readers, even when the text's purpose is to deceive NLP systems or bypass filters. In this research, we expand the study of human suspiciousness by analyzing how individuals perceive adversarial texts. We gather and publish a novel dataset of Likert-scale human evaluations on the suspiciousness of adversarial sentences, crafted by four widely used adversarial attack methods and assess their correlation with the human ability to detect machine-generated alterations. Additionally, we develop a regression-based model to quantify suspiciousness and establish a baseline for future research in reducing the suspiciousness in adversarial text generation. We also demonstrate how the regressor-generated suspicious scores can be incorporated into adversarial generation methods to produce texts that are less likely to be perceived as computer-generated. We make our human suspiciousness annotated data and our code available.

对抗文本可疑度人类评估生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。