arXiv:2410.17439cs.CLcs.AI2024-10被引 3

研究AI生成作文的特征,揭示自动评分与学术诚信的挑战

AI-generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity

  • 用大规模数据对比主流大模型生成作文的特点
  • 发现现有自动评分系统对AI作文识别能力不足
  • 证明跨模型检测可行,有助于维护学术诚信

大型语言模型(LLMs)的快速发展使得生成连贯文章成为可能,推动了教育和职业场景中AI辅助写作的普及。本文基于大规模实证数据,分析并基准测试了主流大模型生成文章的特征与质量,探讨其对写作评估两大核心环节——自动评分与学术诚信——的影响。研究发现,现有自动评分系统(如e-rater)在应用于由AI生成或受其显著影响的文章时存在局限性,需引入新特征以捕捉深层思维,并重新校准特征权重。尽管担心多种大模型的涌现可能使检测变得不可行,但结果显示,针对某一模型训练的检测器通常能以高精度识别其他模型生成的内容,表明实际检测仍具可行性。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has enabled the generation of coherent essays, making AI-assisted writing increasingly common in educational and professional settings. Using large-scale empirical data, we examine and benchmark the characteristics and quality of essays generated by popular LLMs and discuss their implications for two key components of writing assessments: automated scoring and academic integrity. Our findings highlight limitations in existing automated scoring systems, such as e-rater, when applied to essays generated or heavily influenced by AI, and identify areas for improvement, including the development of new features to capture deeper thinking and recalibrating feature weights. Despite growing concerns that the increasing variety of LLMs may undermine the feasibility of detecting AI-generated essays, our results show that detectors trained on essays generated from one model can often identify texts from others with high accuracy, suggesting that effective detection could remain manageable in practice.

AI写作自动评分学术诚信大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。