大模型自检生成内容效果因任务而异,编程题准但短答题常误判。
Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content

- 用不同提示和格式测试大模型识别自己生成内容的能力。
- 编程和长反思题检测准确率高,短问答题反而更像真人。
- 提示设计和回答长度显著影响检测效果,适合教育场景评估。
随着大型语言模型(LLMs)被学生用于生成自然语言回答和程序代码,越来越多关注其是否能自我识别生成内容。本文研究了在编程练习、反思性写作和简答题等多种教育任务中,大模型识别自身输出的能力。基于真实学生作答和多种变体的LLM生成答案,我们评估了不同提示策略与输出格式下的检测表现。研究涵盖三个问题:(1)大模型在不同任务领域识别自身输出的准确性;(2)提示设计、回答长度和任务类型对检测效果的影响;(3)生成内容的哪些特征导致成功或失败检测。结果表明,大模型自检能力具有明显任务依赖性:编程任务和较长反思文本检测较可靠,但在短问答题中,模型常将自身输出判断为比真实学生作答更“人类化”。进一步发现,提示框架和回答冗长度对反思写作检测影响显著,微小提示变化即大幅降低准确率,而编程相关检测对提示变动更具鲁棒性。这些结果揭示了大模型自检在教育场景中的潜力与局限,提示应谨慎将其作为独立工具识别学生AI作业。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly used by students to generate natural language responses and program code, there is growing interest in whether LLMs themselves can be used to distinguish AI-generated work from human-authored submissions. In this paper, we investigate the extent to which LLMs can detect their own generated content across multiple educational task types, including programming exercises, reflective writing, and short-answer questions. Using authentic student responses and multiple variants of LLM-generated answers, we evaluate detection performance under different prompting strategies and output formats. Our study addresses three research questions: (1) how accurately LLMs can identify their own outputs across task domains, (2) how detection effectiveness is influenced by factors such as prompt design, response length, and task type, and (3) what characteristics of LLM-generated responses contribute to successful or failed detection. Our findings show that LLM-based detection is highly task-dependent: detection is substantially more reliable for programming tasks and longer reflective responses, but performs poorly for short-answer questions, where LLMs frequently judge their own outputs as more human-like than authentic student responses. We further find that prompt framing and response verbosity have a pronounced effect on detectability in reflective writing tasks, with relatively minor prompt variations significantly reducing detection accuracy, while programming-related detection is more robust to prompt changes. Together, these results highlight both the potential and the limitations of LLM self-detection in educational settings and suggest caution in relying on LLMs as standalone tools for identifying AI-generated student work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。