提出可有效识别经人工化工具篡改的AI文本的检测模型。
DAMAGE: Detecting Adversarially Modified AI Generated Text
- 采用数据驱动增强方法构建鲁棒检测模型
- 在19种人类化工具上保持高检测率且误报率低
- 对自训练对抗模型仍具泛化能力,适合安全防护场景
AI人类化工具是一类新型在线软件,旨在改写AI生成文本以逃避检测。我们研究了19种AI人类化与重写工具,定性评估其对原文语义的保真度及影响。结果显示,许多现有AI检测器无法识别经人类化处理的文本。为此,我们提出一种基于数据中心增强的鲁棒检测模型,可在多种人类化工具间保持高检测准确率,并维持低误报率。我们还主动攻击自身检测器,训练对抗模型以优化其规避预测,结果表明该检测器在跨工具场景下仍具备足够鲁棒性。
原文摘要 · Abstract (English)
AI humanizers are a new class of online software tools meant to paraphrase and rewrite AI-generated text in a way that allows them to evade AI detection software. We study 19 AI humanizer and paraphrasing tools and qualitatively assess their effects and faithfulness in preserving the meaning of the original text. We show that many existing AI detectors fail to detect humanized text. Finally, we demonstrate a robust model that can detect humanized AI text while maintaining a low false positive rate using a data-centric augmentation approach. We attack our own detector, training our own fine-tuned model optimized against our detector's predictions, and show that our detector's cross-humanizer generalization is sufficient to remain robust to this attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。