arXiv:2509.20691cs.CL2025-09EMNLP

提出新型红鲱鱼攻击,让检测模型误判但分类仍对,暴露检测可靠性隐患。

RedHerring Attack: Testing the Reliability of Attack Detection

  • 设计红鲱鱼攻击:改写文本使检测器误报,分类器仍正确
  • 在4个数据集上使检测准确率下降20至71点,分类性能不变或提升
  • 适合关注AI安全与检测系统可信度的研究者

针对对抗性文本攻击,已有攻击检测模型被提出并成功识别篡改文本。这些检测模型可作为NLP模型的额外校验,并为人工干预提供信号。然而,其可靠性尚未得到充分探索。为此,我们提出并测试了一种新型攻击设定——红鲱鱼(RedHerring)攻击:通过修改文本,使检测模型误判为存在攻击,同时保持分类器判断正确,从而制造分类器与检测器之间的矛盾。若人类发现检测器发出错误预警而分类器输出正确,将质疑检测器的可靠性。我们在4个数据集上对3种检测器(保护4种分类器)进行了测试,结果表明红鲱鱼攻击可导致检测准确率下降20至71个百分点,同时维持(或提升)分类器准确率。作为初步防御,我们提出一种无需重新训练分类器或检测器的简单置信度检查机制,显著提升检测准确率。该威胁模型揭示了对手可能针对检测系统的新攻击路径。

原文摘要 · Abstract (English)

In response to adversarial text attacks, attack detection models have been proposed and shown to successfully identify text modified by adversaries. Attack detection models can be leveraged to provide an additional check for NLP models and give signals for human input. However, the reliability of these models has not yet been thoroughly explored. Thus, we propose and test a novel attack setting and attack, RedHerring. RedHerring aims to make attack detection models unreliable by modifying a text to cause the detection model to predict an attack, while keeping the classifier correct. This creates a tension between the classifier and detector. If a human sees that the detector is giving an ``incorrect'' prediction, but the classifier a correct one, then the human will see the detector as unreliable. We test this novel threat model on 4 datasets against 3 detectors defending 4 classifiers. We find that RedHerring is able to drop detection accuracy between 20 - 71 points, while maintaining (or improving) classifier accuracy. As an initial defense, we propose a simple confidence check which requires no retraining of the classifier or detector and increases detection accuracy greatly. This novel threat model offers new insights into how adversaries may target detection models.

对抗攻击检测可靠性NLP安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。