arXiv:2508.08096cs.CLcs.LG2025-08被引 7

测试大模型文本检测在教育场景下的表现,发现学生参与度影响检测准确性。

Assessing LLM Text Detection in Educational Contexts: Does Human Contribution Affect Detection?

  • 提出学生贡献度分级,涵盖从纯人工到完全生成再到伪装生成的五类文本
  • 超过900篇学生作文与12500篇生成文构成新数据集,揭示中间层级文本检测率不足60%
  • 检测器易误判人类修改稿为机器生成,对教育公平性构成潜在风险

大语言模型(LLM)的发展和普及使得学生能轻易自动生成文本,给教育机构带来学术诚信挑战。为此,自动检测生成文本的学习分析方法日益受到关注。本文在教育场景下基准测试多种前沿检测器性能,构建了名为GEDE的新数据集,包含超过900篇学生手写作文和超过12500篇跨领域生成文本。为捕捉实际中多样化的使用模式,提出“贡献水平”概念,涵盖纯人工写作、轻微修改的生成文本、完全生成文本,以及通过“人性化”策略主动攻击检测器的文本。结果显示,大多数检测器在中等贡献度文本(如经生成模型优化的人工写作)上分类准确率显著下降,尤其容易产生误报,这在教育环境中可能严重损害学生权益。相关数据集、代码及补充材料已公开于https://github.com/lukasgehring/Assessing-LLM-Text-Detection-in-Educational-Contexts。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) and their increased accessibility have made it easier than ever for students to automatically generate texts, posing new challenges for educational institutions. To enforce norms of academic integrity and ensure students' learning, learning analytics methods to automatically detect LLM-generated text appear increasingly appealing. This paper benchmarks the performance of different state-of-the-art detectors in educational contexts, introducing a novel dataset, called Generative Essay Detection in Education (GEDE), containing over 900 student-written essays and over 12,500 LLM-generated essays from various domains. To capture the diversity of LLM usage practices in generating text, we propose the concept of contribution levels, representing students' contribution to a given assignment. These levels range from purely human-written texts, to slightly LLM-improved versions, to fully LLM-generated texts, and finally to active attacks on the detector by "humanizing" generated texts. We show that most detectors struggle to accurately classify texts of intermediate student contribution levels, like LLM-improved human-written texts. Detectors are particularly likely to produce false positives, which is problematic in educational settings where false suspicions can severely impact students' lives. Our dataset, code, and additional supplementary materials are publicly available at https://github.com/lukasgehring/Assessing-LLM-Text-Detection-in-Educational-Contexts.

文本检测教育AILLM安全误报风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。