用AI漏洞反制AI作弊,通过视觉扰动识别学生是否抄袭
Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating

- 用细微视觉扰动让AI答错,制造可检测的答题模式
- 在三款主流AI模型上验证,扰动后错误答案出现率超70%
- 适合教育机构用于检测作业抄袭,尤其针对图文题
生成式AI的普及使学生可将认知任务外包给日益强大的助手,看似能力提升,实则削弱独立思考能力。本文探究能否利用对抗性机器学习保护教育练习免受此类依赖影响。方法为设计多模态选择题,其图像部分加入微小扰动,诱导AI解题器持续选择特定错误答案。这些错误答案形成统计指纹:盲目复制AI结果的学生重复该模式的频率远高于真实作答者。我们在三种主流多模态大模型(Anthropic Claude、Google Gemini、OpenAI ChatGPT)上验证该方案可行性,使用可获取的代理模型优化扰动,确保响应模式一致。通过统计假设检验可实现可靠检测。研究证明,以机器漏洞对抗机器辅助推理具有潜力,但也存在局限性。
原文摘要 · Abstract (English)
The widespread adoption of generative AI enables students to outsource cognitive effort to increasingly capable assistants, creating an illusion of competence while undermining the independent reasoning that education aims to cultivate. We investigate whether adversarial machine learning can be repurposed to protect educational exercises against such corrosive reliance. Our approach uses multimodal multiple-choice questions whose visual components can be protected with subtle visual perturbations that steer AI solvers toward designated incorrect answers. These responses form a statistical fingerprint: students who blindly copy a solver reproduce the induced answer pattern more frequently than genuine students. We study the feasibility of this paradigm under realistic black-box assistant assumptions using three of the most common state-of-the-art multimodal language models: Anthropic's Claude, Google's Gemini, and OpenAI's ChatGPT. By using accessible surrogate models, we optimize adversarial perturbations that induce consistent response patterns. Those patterns enable principled detection through statistical hypothesis testing. These findings establish both the promise and the limitations of fighting machine-assisted reasoning with the vulnerabilities of the machines themselves.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。