用相互关联的题目设计考试,能有效防止AI代写,更真实评估学生能力。
Designing AI-Resilient Assessments Using Interconnected Problems: A Theoretically Grounded and Empirically Validated Framework
- 把题目串成链条,后一题依赖前一题输出,迫使AI无法跳过推理。
- 用固定标准的半结构化题,学生得分比开放题高30个百分点,且更难被AI作弊。
- 适合教育工作者参考,尤其在数据科学、编程等需要真实能力评估的课程。
生成式AI工具的普及使传统模块化考核在计算机与数据类教育中日益失效,导致学术评价与真实能力测量脱节。本文提出一个理论扎实且经实证验证的AI抗性评估设计框架。首先,建立两个形式化命题:(1)由相互关联的问题组成的评估——后续任务依赖前序任务输出——因依赖多步推理和持续上下文,天然比模块化评估更具抗AI能力;(2)具有确定性评判标准的半结构化问题,比完全开放的项目更能可靠衡量学生能力,后者易被AI套用惯用模板。其次,基于三所大学数据科学课程(N=117)的实证分析验证上述命题:学生使用AI完成模块化作业几乎全对,但在监考考试中表现下降约30个百分点(Cohen d = 1.51)。相比之下,相互关联项目与模块化测试高度一致(r = 0.954, p < 0.001),同时保持抗AI性;而监考考试一致性较弱(r = 0.726, p < 0.001)。最后,将研究成果转化为可操作的评估设计流程,帮助教师构建促进深度参与、贴近行业实践且能抵御简单AI代劳的考核方式。
原文摘要 · Abstract (English)
The proliferation of generative AI tools has rendered traditional modular assessments in computing and data-centric education increasingly ineffective, creating a disconnect between academic evaluation and authentic skill measurement. This paper presents a theoretically grounded framework for designing AI-resilient assessments, supported by formal analysis and empirical validation. We make three primary contributions. First, we establish two formal propositions. (1) Assessments composed of interconnected problems, in which outputs serve as inputs to subsequent tasks, are inherently more AI-resilient than modular assessments due to their reliance on multi-step reasoning and sustained context. (2) Semi-structured problems with deterministic success criteria provide more reliable measures of student competency than fully open-ended projects, which allow AI systems to default to familiar solution templates. These results challenge widely cited recommendations in recent institutional and policy guidance that promote open-ended assessments as inherently more robust to AI assistance. Second, we validate these propositions through empirical analysis of three university data science courses (N = 117). We observe a substantial AI inflation effect: students achieve near-perfect scores on AI-assisted modular homework, while performance drops by approximately 30 percentage points on proctored exams (Cohen d = 1.51). In contrast, interconnected projects remain strongly aligned with modular assessments (r = 0.954, p < 0.001) while maintaining AI resistance, whereas proctored exams show weaker alignment (r = 0.726, p < 0.001). Third, we translate these findings into a practical assessment design procedure that enables educators to construct evaluations that promote deeper engagement, reflect industry practice, and resist trivial AI delegation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。