arXiv:2410.05573cs.CRcs.AI2024-10NAACL被引 2

构建高质量有毒对抗文本数据集,提升内容审核模型鲁棒性

TaeBench: Improving Quality of Toxic Adversarial Examples

  • 设计自动化标注+人工验证的双轨质量控制流程
  • 筛选出26.4万条语义自然且能骗过检测模型的对抗样本
  • 适合研究对抗攻击与内容安全防护的学者使用

毒性文本检测系统可能受到对抗样本的威胁——对输入文本进行微小扰动即可误导系统做出错误判断。现有攻击算法耗时且常生成无效或模糊的对抗样本,难以用于真实场景的内容审核评估与优化。本文提出一种针对有毒对抗文本(TAE)的质量控制注释流程,结合基于模型的自动标注与人工质量验证,确保生成的TAE满足:欺骗目标毒性模型产生良性预测、语法合理、风格自然、具备语义毒性。在超过20个当前最优(SOTA)TAE攻击方法生成的94万条原始样本中,发现大量无效样本。通过该流程筛选并整理出高质量数据集TaeBench(共26.4万条)。实证表明,TaeBench可有效攻击SOTA毒性内容审核模型与服务;同时,利用TaeBench进行对抗训练,显著提升了两个毒性检测器的鲁棒性。

原文摘要 · Abstract (English)

Toxicity text detectors can be vulnerable to adversarial examples - small perturbations to input text that fool the systems into wrong detection. Existing attack algorithms are time-consuming and often produce invalid or ambiguous adversarial examples, making them less useful for evaluating or improving real-world toxicity content moderators. This paper proposes an annotation pipeline for quality control of generated toxic adversarial examples (TAE). We design model-based automated annotation and human-based quality verification to assess the quality requirements of TAE. Successful TAE should fool a target toxicity model into making benign predictions, be grammatically reasonable, appear natural like human-generated text, and exhibit semantic toxicity. When applying these requirements to more than 20 state-of-the-art (SOTA) TAE attack recipes, we find many invalid samples from a total of 940k raw TAE attack generations. We then utilize the proposed pipeline to filter and curate a high-quality TAE dataset we call TaeBench (of size 264k). Empirically, we demonstrate that TaeBench can effectively transfer-attack SOTA toxicity content moderation models and services. Our experiments also show that TaeBench with adversarial training achieve significant improvements of the robustness of two toxicity detectors.

对抗样本内容审核数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。