arXiv:2412.02713cs.AI2024-12中稿 · The 15th Internati…被引 6

用答题模式差异识别AI作弊,专治多选题代考

Applying IRT to Distinguish Between Human and Generative AI Responses to Multiple-Choice Assessments

  • 基于项目反应理论分析答题模式差异
  • 能有效区分人类与主流大模型的作答行为
  • 适合教育评估、考试防作弊场景使用

生成式AI正深刻改变教育领域,引发作弊担忧。尽管多选题广泛用于测评,但针对此类题型的AI作弊检测几乎未受关注,远落后于对文本类作业中AI作弊的研究。本文提出一种基于项目反应理论(IRT)的方法,假设人工与人工智能在答题模式上存在本质差异,AI作弊会表现为对人类预期响应模式的偏离,通过个体拟合统计量进行建模。实验表明,该方法能有效区分人类与顶级聊天机器人(ChatGPT、Claude、Gemini)生成的答案,且对作弊程度敏感;同时揭示不同模型具有不同的推理特征。本研究为在多选题测评中识别AI作弊提供了理论基础与实证支持。

原文摘要 · Abstract (English)

Generative AI is transforming the educational landscape, raising significant concerns about cheating. Despite the widespread use of multiple-choice questions in assessments, the detection of AI cheating in MCQ-based tests has been almost unexplored, in contrast to the focus on detecting AI-cheating on text-rich student outputs. In this paper, we propose a method based on the application of Item Response Theory to address this gap. Our approach operates on the assumption that artificial and human intelligence exhibit different response patterns, with AI cheating manifesting as deviations from the expected patterns of human responses. These deviations are modeled using Person-Fit Statistics. We demonstrate that this method effectively highlights the differences between human responses and those generated by premium versions of leading chatbots (ChatGPT, Claude, and Gemini), but that it is also sensitive to the amount of AI cheating in the data. Furthermore, we show that the chatbots differ in their reasoning profiles. Our work provides both a theoretical foundation and empirical evidence for the application of IRT to identify AI cheating in MCQ-based assessments.

AI检测教育评测项目反应理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。