arXiv:2410.03658cs.CLcs.LG2024-10EMNLP被引 13

提出无语法错误的黑盒攻击,可骗过所有文本检测器

RAFT: Realistic Attacks to Fool Text Detectors

论文配图:RAFT: Realistic Attacks to Fool Text Detectors
图 1 · 摘自论文原文
  • 利用词级嵌入迁移性,精准选择需扰动词汇
  • 攻击成功率最高达99%,且跨模型有效
  • 生成内容逼真,适合训练抗攻击检测器

大型语言模型(LLMs)在各类任务中表现出卓越的流畅性,但其被用于传播虚假信息等不道德行为的问题日益严重。尽管已有多种LLM检测方法被提出,其鲁棒性和可靠性仍不明确。本文提出RAFT:一种无语法错误的黑盒攻击方法,针对现有LLM检测器。与以往针对语言模型的攻击不同,该方法利用词级嵌入的可迁移性,在保持原文质量的前提下进行扰动。通过辅助嵌入贪婪选择待扰动词汇。实验表明,该攻击在多个领域内对所有检测器的有效性最高可达99%,且具备跨源模型的迁移能力。人工评估显示,攻击生成的内容与原始人类撰写文本难以区分。此外,RAFT生成样本可用于训练对抗鲁棒的检测器。本工作揭示当前LLM检测器缺乏对抗鲁棒性,凸显亟需更稳健的检测机制。

原文摘要 · Abstract (English)

Large language models (LLMs) have exhibited remarkable fluency across various tasks. However, their unethical applications, such as disseminating disinformation, have become a growing concern. Although recent works have proposed a number of LLM detection methods, their robustness and reliability remain unclear. In this paper, we present RAFT: a grammar error-free black-box attack against existing LLM detectors. In contrast to previous attacks for language models, our method exploits the transferability of LLM embeddings at the word-level while preserving the original text quality. We leverage an auxiliary embedding to greedily select candidate words to perturb against the target detector. Experiments reveal that our attack effectively compromises all detectors in the study across various domains by up to 99%, and are transferable across source models. Manual human evaluation studies show our attacks are realistic and indistinguishable from original human-written text. We also show that examples generated by RAFT can be used to train adversarially robust detectors. Our work shows that current LLM detectors are not adversarially robust, underscoring the urgent need for more resilient detection mechanisms.

文本检测对抗攻击大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。