测试6种AI检测工具对DeepSeek生成文本的识别能力,发现人类化改写可显著降低检测准确率。
Evaluating the Performance of AI Text Detectors, Few-Shot and Chain-of-Thought Prompting Using DeepSeek Generated Text
- 用深度思考和少样本提示让DeepSeek自己当检测器,效果优于多数现成工具。
- 人类化改写使主流检测器准确率降至52%~71%,威胁检测可靠性。
- 少样本提示下检测准确率达96%以上,适合需要高精度识别的场景。
大型语言模型(LLM)迅速改变了文本创作方式,引发对写作真实性的担忧,推动了人工智能(AI)检测技术的发展。对抗性攻击如标准改写和人类化改写会削弱检测器识别机器生成文本的能力。以往研究主要针对ChatGPT等知名模型,但对近期发布的DeepSeek模型缺乏系统评估。本文考察六种常见AI检测工具——AI Text Classifier、Content Detector AI、Copyleaks、QuillBot、GPT-2和GPTZero——对DeepSeek-v3生成文本的识别能力,涵盖原始文本及经过改写与人类化处理的样本。研究收集49组人工撰写的问答对,并由DeepSeek-v3生成对应回答,再通过对抗技术生成196个新样本以测试检测器鲁棒性。结果显示,QuillBot与Copyleaks在原始及改写文本上表现接近完美,而AI Text Classifier和GPT-2则结果不稳定。最有效的攻击为人类化改写,使Copyleaks准确率降至71%,QuillBot降至58%,GPTZero降至52%。采用少样本提示与链式思维(CoT)推理时,检测准确率显著提升,最佳五样本提示仅误判1例(AI召回率96%,人类召回率100%)。
原文摘要 · Abstract (English)
Large language models (LLMs) have rapidly transformed the creation of written materials. LLMs have led to questions about writing integrity, thereby driving the creation of artificial intelligence (AI) detection technologies. Adversarial attacks, such as standard and humanized paraphrasing, inhibit detectors' ability to detect machine-generated text. Previous studies have mainly focused on ChatGPT and other well-known LLMs and have shown varying accuracy across detectors. However, there is a clear gap in the literature about DeepSeek, a recently published LLM. Therefore, in this work, we investigate whether six generally accessible AI detection tools -- AI Text Classifier, Content Detector AI, Copyleaks, QuillBot, GPT-2, and GPTZero -- can consistently recognize text generated by DeepSeek. The detectors were exposed to the aforementioned adversarial attacks. We also considered DeepSeek as a detector by performing few-shot prompting and chain-of-thought reasoning (CoT) for classifying AI and human-written text. We collected 49 human-authored question-answer pairs from before the LLM era and generated matching responses using DeepSeek-v3, producing 49 AI-generated samples. Then, we applied adversarial techniques such as paraphrasing and humanizing to add 196 more samples. These were used to challenge detector robustness and assess accuracy impact. While QuillBot and Copyleaks showed near-perfect performance on original and paraphrased DeepSeek text, others -- particularly AI Text Classifier and GPT-2 -- showed inconsistent results. The most effective attack was humanization, reducing accuracy to 71% for Copyleaks, 58% for QuillBot, and 52% for GPTZero. Few-shot and CoT prompting showed high accuracy, with the best five-shot result misclassifying only one of 49 samples (AI recall 96%, human recall 100%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。