动态上下文扰动让文本攻击更隐蔽且有效
Adversarial Text Generation with Dynamic Contextual Perturbation
- 基于预训练模型动态生成跨句段的语境感知扰动
- 在多个数据集上成功欺骗主流NLP模型,攻击成功率更高
- 适合研究模型鲁棒性或防御机制的学者参考
自然语言处理模型的对抗攻击通过在输入文本中引入细微扰动暴露其脆弱性,通常在保持人类可读性的前提下导致误分类。现有方法多聚焦于词级或局部片段修改,忽略整体上下文,导致扰动可检测或语义不一致。本文提出一种新型对抗文本攻击方法——动态上下文扰动(DCP),该方法在句子、段落和文档层面动态生成语境感知的扰动,确保语义一致性和语言流畅性。利用预训练语言模型能力,DCP通过对抗目标函数迭代优化扰动,在诱导模型误分类与保持文本自然性之间取得平衡。实验结果表明,DCP在多种NLP模型和数据集上均显著提升了攻击效果,生成的对抗样本更贴近自然语言模式。本研究强调了上下文在对抗攻击中的关键作用,为构建更具鲁棒性的NLP系统提供了基础。
原文摘要 · Abstract (English)
Adversarial attacks on Natural Language Processing (NLP) models expose vulnerabilities by introducing subtle perturbations to input text, often leading to misclassification while maintaining human readability. Existing methods typically focus on word-level or local text segment alterations, overlooking the broader context, which results in detectable or semantically inconsistent perturbations. We propose a novel adversarial text attack scheme named Dynamic Contextual Perturbation (DCP). DCP dynamically generates context-aware perturbations across sentences, paragraphs, and documents, ensuring semantic fidelity and fluency. Leveraging the capabilities of pre-trained language models, DCP iteratively refines perturbations through an adversarial objective function that balances the dual objectives of inducing model misclassification and preserving the naturalness of the text. This comprehensive approach allows DCP to produce more sophisticated and effective adversarial examples that better mimic natural language patterns. Our experimental results, conducted on various NLP models and datasets, demonstrate the efficacy of DCP in challenging the robustness of state-of-the-art NLP systems. By integrating dynamic contextual analysis, DCP significantly enhances the subtlety and impact of adversarial attacks. This study highlights the critical role of context in adversarial attacks and lays the groundwork for creating more robust NLP systems capable of withstanding sophisticated adversarial strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。