arXiv:2601.12505cs.CL2026-01被引 1

用语义陷阱干扰AI答题,保护学术考试安全

DoPE: Decoy Oriented Perturbation Encapsulation Human-Readable, AI-Hostile Documents for Academic Integrity

  • 在文档中嵌入语义陷阱,利用AI解析差异制造干扰
  • 96.3%的攻击尝试被阻止或诱导出错,检测率达91.4%
  • 适合考试系统开发者和学术诚信研究者使用

多模态大语言模型可直接处理考试文档,威胁传统评估与学术诚信。本文提出DoPE(诱饵导向扰动封装)框架,在文档层嵌入语义诱饵,利用MLLM处理流程中的渲染-解析差异实现防御。通过作者端注入,DoPE提供模型无关的防解题(阻止或混淆自动作答)与检测(标记盲目依赖AI行为),无需依赖传统的一次性分类器。我们形式化了预防与检测任务,提出基于LLM引导的FewSoRT-Q生成题级语义诱饵,并通过FewSoRT-D将其封装为带水印文档。在新构建的Integrity-Bench基准上测试,该基准包含1826份来自公开QA数据集与OpenCourseWare的PDF+HTML考试文件。针对OpenAI与Anthropic的黑盒MLLM,DoPE实现91.4%检测率(误报率8.7%),并在96.3%的攻击尝试中成功阻止完成或诱导诱饵对齐错误。我们开源Integrity-Bench、工具包及评估代码,支持学术诚信文档层防御的可复现研究。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) can directly consume exam documents, threatening conventional assessments and academic integrity. We present DoPE (Decoy-Oriented Perturbation Encapsulation), a document-layer defense framework that embeds semantic decoys into PDF/HTML assessments to exploit render-parse discrepancies in MLLM pipelines. By instrumenting exams at authoring time, DoPE provides model-agnostic prevention (stop or confound automated solving) and detection (flag blind AI reliance) without relying on conventional one-shot classifiers. We formalize prevention and detection tasks, and introduce FewSoRT-Q, an LLM-guided pipeline that generates question-level semantic decoys and FewSoRT-D to encapsulate them into watermarked documents. We evaluate on Integrity-Bench, a novel benchmark of 1826 exams (PDF+HTML) derived from public QA datasets and OpenCourseWare. Against black-box MLLMs from OpenAI and Anthropic, DoPE yields strong empirical gains: a 91.4% detection rate at an 8.7% false-positive rate using an LLM-as-Judge verifier, and prevents successful completion or induces decoy-aligned failures in 96.3% of attempts. We release Integrity-Bench, our toolkit, and evaluation code to enable reproducible study of document-layer defenses for academic integrity.

学术诚信AI防御文档安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。