arXiv:2506.23517cs.AIcs.CL2025-06被引 3

测试GPTZero识别AI写作的准确率,发现长文本检测效果好但易误判人类作文。

Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays

  • 用长短不一的随机作文测试GPTZero,分短(40-100字)、中(100-350字)、长(350-800字)三类
  • 纯AI作文91%-100%被正确识别,但部分人类作文被误判为AI生成
  • 提示教师勿仅依赖工具,需结合人工判断

随着学生使用AI工具日益普遍,教师开始采用GPTZero、QuillBot等AI检测工具识别文本来源。本研究聚焦于最常用的GPTZero在识别不同长度随机提交作文中的表现:短篇(40-100字)、中篇(100-350字)、长篇(350-800字)。我们收集了28篇AI生成论文和50篇人类撰写论文,逐篇输入GPTZero,记录其生成概率与置信度。结果显示,绝大多数AI生成论文被准确识别(91%-100%判定为AI生成),而人类论文则出现波动,存在少量误判情况。这表明GPTZero对纯AI内容识别有效,但在区分人类写作方面可靠性有限。教育工作者应谨慎依赖此类工具。

原文摘要 · Abstract (English)

As the use of AI tools by students has become more prevalent, instructors have started using AI detection tools like GPTZero and QuillBot to detect AI written text. However, the reliability of these detectors remains uncertain. In our study, we focused mostly on the success rate of GPTZero, the most-used AI detector, in identifying AI-generated texts based on different lengths of randomly submitted essays: short (40-100 word count), medium (100-350 word count), and long (350-800 word count). We gathered a data set consisting of twenty-eight AI-generated papers and fifty human-written papers. With this randomized essay data, papers were individually plugged into GPTZero and measured for percentage of AI generation and confidence. A vast majority of the AI-generated papers were detected accurately (ranging from 91-100% AI believed generation), while the human generated essays fluctuated; there were a handful of false positives. These findings suggest that although GPTZero is effective at detecting purely AI-generated content, its reliability in distinguishing human-authored texts is limited. Educators should therefore exercise caution when relying solely on AI detection tools.

AI检测GPTZero教育应用误判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。