测试不同AI模型生成作文的检测效果,帮教育评估更可信
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
- 用公开GRE题目的作文数据测试检测器跨模型泛化能力
- 发现训练用某模型的检测器对其他模型生成文识别率下降
- 为教育场景下检测工具的选用和更新提供实证建议
写作是基础读写能力的核心,支撑有效沟通、批判性思维、跨学科学习以及复杂思想的组织表达。因此,写作评估在衡量语言水平、交流效果和分析推理方面至关重要。大型语言模型(LLMs)的快速发展使得生成连贯高质量的作文变得极为容易,引发了对学生提交作品真实性的严重担忧。本文首先概述当前针对AI生成与辅助作文的检测技术现状,并提出负责任使用指南。随后,基于对公共GRE写作题目的响应作文,通过实证分析评估了以某一LLM训练的检测器在识别其他LLM生成作文时的泛化性能。研究结果为检测器的实际应用开发与再训练提供了重要指导。
原文摘要 · Abstract (English)
Writing is a foundational literacy skill that underpins effective communication, fosters critical thinking, facilitates learning across disciplines, and enables individuals to organize and articulate complex ideas. Consequently, writing assessment plays a vital role in evaluating language proficiency, communicative effectiveness, and analytical reasoning. The rapid advancement of large language models (LLMs) has made it increasingly easy to generate coherent, high-quality essays, raising significant concerns about the authenticity of student-submitted work. This chapter first provides an overview of the current landscape of detectors for AI-generated and AI-assisted essays, along with guidelines for their responsible use. It then presents empirical analyses to evaluate how well detectors trained on essays from one LLM generalize to identifying essays produced by other LLMs, based on essays generated in response to public GRE writing prompts. These findings provide guidance for developing and retraining detectors for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。