arXiv:2606.30655cs.CYcs.AI2026-06

用帕累托盈余机制评估学生是否超越AI基础能力,实现公平的智能时代考核。

Toward AI-Resilient Assessment in Computer Science Courses in an AI-Native World

  • 以帕累托盈余为评分依据,衡量学生成果是否优于预设AI基准
  • 在允许自由使用AI前提下,避免私有资源或频繁调用带来不公平优势
  • 适用于高年级计算机课程,适合关注真实创新能力培养的教学场景

在人工智能原生的教育环境中,高级计算机科学课程的评估应聚焦于‘抗AI技能’:即超越强基准AI所能达成成果的能力。此类评估应允许学生自由使用AI,同时降低个人拥有更强私有AI资源或更密集使用AI本身带来的评分优势。本文提出一个最小化的形式化框架,包含具体任务、可执行评估器、声明的AI原生帕累托前沿及基于帕累托盈余的评分规则。核心观点为:帕累托盈余提供了一种可度量、依赖协议的证明,表明提交成果实现了当前声明的AI基线尚未覆盖的权衡;以该盈余为依据进行评分,可相对于该基线实现抗AI性。将盈余解释为学生能力证据需依赖评估协议(如设计报告、消融实验、提示追踪、口头核查或可复现性说明),但评分本身是行为且可执行的。框架进一步扩展至自提升AI循环、预算中立性、服务器中介反馈及基于提示的红队测试等实际挑战。作为实例,本文描述了在莱斯大学COMP 480/580课程中,围绕布隆过滤器设计的抗AI近似成员查询作业,旨在检验学生能否改进超出AI生成实现的表现。

原文摘要 · Abstract (English)

AI-native course assessments in senior computer science courses and related fields should grade students by \emph{AI-resilient skill}: the ability to achieve outcomes beyond a strong AI baseline. Such assessments should allow students to use AI freely, while reducing the extent to which greater private AI budget or more intensive AI use, by itself, becomes a grading advantage. This paper proposes a minimal formal framework for this goal. The framework specifies a real task, an executable evaluator, a declared AI-native Pareto frontier, and a grading rule based on Pareto surplus. The central claim is simple: Pareto surplus provides a measurable, protocol-relative certificate that a submitted artifact achieves a tradeoff not already supplied by the declared AI baseline, and grading by this surplus is AI-resilient with respect to that baseline. Interpreting surplus as evidence of student skill requires the surrounding assessment protocol--for example, design reports, ablations, prompt traces, oral checks, or reproducibility explanations--but the grading certificate itself is behavioral and executable. The framework is then extended to practical complications, including self-improving AI loops, budget neutrality, server-mediated feedback, and prompt-based red teaming. As a concrete instantiation, we describe an AI-resilient approximate-membership assignment centered on Bloom filters for COMP 480/580 at Rice University, designed to test whether students can improve beyond AI-generated implementations.

AI评估抗AI能力教学改革帕累托优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。