提出三维检测框架,提升文本生成内容识别的抗攻击能力。
Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks

- 从内容与表达两方面补充统计检测视角
- 在17种攻击下,准确率提升22.3%(AUROC)
- 适合需要高可靠性的内容安全场景
现有AI生成文本检测方法易受操控文本特征的攻击。本文提出三重视角检测框架Triospect,通过引入文本的核心思想(内容)和风格特征(表达)两个额外维度,在包含17种攻击、12个领域和17个源模型的两个基准上进行实验。结果表明,Triospect对各类攻击具有强鲁棒性:在Humanize-16K后攻击子集上,相比强基线,AUROC提升22.3%,TPR01提升13%;在对抗性RAID数据集上,AUROC提升9.1%,TPR01提升22%。该框架是统计方法在提升检测可靠性方面的开创性尝试。代码与数据已开源:https://github.com/baoguangsheng/triospect。
原文摘要 · Abstract (English)
Existing AI-generated text detectors are vulnerable to attacks that manipulate textual characteristics. In this study, we propose a novel Triospect Detection Framework by using additional perspectives of content (core ideas) and expression (stylistic elements) within a given text. Experiments on two benchmarks involving 17 attacks, 12 domains, and 17 source models demonstrate that Triospect is robust against these attacks. It improves the strong baseline by a significant margin of 22.3% (AUROC) and 13% (TPR01) on the Humanize-16K after-attack subset, and by 9.1% (AUROC) and 22% (TPR01) on the adversarial RAID. This framework marks a pioneering effort in statistical methods to enhance detection reliability against attacks. We release our data and code at https://github.com/baoguangsheng/triospect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。