测试主流文本检测器在真实场景下的失效情况,发现轻易可被绕过。
A Practical Examination of AI-Generated Text Detectors for Large Language Models
- 用多种提示策略模拟实际攻击,测试未见过的模型和数据集
- 部分检测器在0.01假阳性率下真阳性率低至0%
- 无论训练或零样本检测器,都难兼顾灵敏度与准确率
大型语言模型的普及引发了对其滥用的担忧,尤其是生成文本被误标为人类撰写的情况。现有文本检测器声称能在多种条件下有效识别此类内容。本文通过在未接触过的领域、数据集和模型上评估多个主流检测器(RADAR、Wild、T5Sentinel、Fast-DetectGPT、PHD、LogRank、Binoculars),并采用多种提示策略模拟实际对抗攻击,发现仅需适度努力即可显著规避检测。研究强调在特定假阳性率(TPR@FPR)下的真正例率的重要性,结果表明这些检测器在某些场景下表现极差,最高假阳性率0.01时真阳性率低至0%。研究显示,无论是训练过的还是零样本检测器,均难以在保持合理假阳性率的同时维持高灵敏度。
原文摘要 · Abstract (English)
The proliferation of large language models has raised growing concerns about their misuse, particularly in cases where AI-generated text is falsely attributed to human authors. Machine-generated content detectors claim to effectively identify such text under various conditions and from any language model. This paper critically evaluates these claims by assessing several popular detectors (RADAR, Wild, T5Sentinel, Fast-DetectGPT, PHD, LogRank, Binoculars) on a range of domains, datasets, and models that these detectors have not previously encountered. We employ various prompting strategies to simulate practical adversarial attacks, demonstrating that even moderate efforts can significantly evade detection. We emphasize the importance of the true positive rate at a specific false positive rate (TPR@FPR) metric and demonstrate that these detectors perform poorly in certain settings, with [email protected] as low as 0%. Our findings suggest that both trained and zero-shot detectors struggle to maintain high sensitivity while achieving a reasonable true positive rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。