arXiv:2503.15772cs.DLcs.AI2025-03被引 31

用隐写水印检测大模型生成的论文评审,准确率高且抗防御。

Detecting LLM-Generated Peer Reviews

  • 通过间接提示注入将水印嵌入生成评审,隐蔽性强。
  • 在多种大模型上水印嵌入成功率超90%,错误率可控。
  • 适合期刊编辑部、科研诚信机构用于识别AI辅助评审。

同行评审的可靠性是科学进步的基础,但大语言模型(LLMs)的兴起引发了担忧:部分评审人可能借助这些工具生成评审意见而非独立撰写。尽管一些期刊已禁止使用LLM辅助评审,但执行困难,因现有检测工具无法可靠区分完全由模型生成的评审与仅经AI润色的评审。本文提出一种检测大模型生成评审的方法:通过论文PDF实施间接提示注入,促使模型在生成评审时嵌入隐蔽水印,随后检测水印是否存在。我们识别并解决了该方法中若干常见陷阱。主要贡献在于构建了一个严格的水印与检测框架,提供强大的统计保障。具体而言,提出了水印方案与假设检验方法,可控制多篇评审中的族误差率(family-wise error rate),在不假设人类写作特征的前提下,比传统校正方法(如Bonferroni)具有更高的统计功效。我们探索了多种间接提示注入策略——包括基于字体的嵌入和混淆提示——并在多种评审防御场景下评估其有效性。实验表明,水印嵌入在多种大模型上成功率高;同时,该方法对常见防御手段具有鲁棒性,且统计测试的误差界在实践中成立。相比之下,基于Bonferroni的校正过于保守,难以实用。

原文摘要 · Abstract (English)

The integrity of peer review is fundamental to scientific progress, but the rise of large language models (LLMs) has introduced concerns that some reviewers may rely on these tools to generate reviews rather than writing them independently. Although some venues have banned LLM-assisted reviewing, enforcement remains difficult as existing detection tools cannot reliably distinguish between fully generated reviews and those merely polished with AI assistance. In this work, we address the challenge of detecting LLM-generated reviews. We consider the approach of performing indirect prompt injection via the paper's PDF, prompting the LLM to embed a covert watermark in the generated review, and subsequently testing for presence of the watermark in the review. We identify and address several pitfalls in naïve implementations of this approach. Our primary contribution is a rigorous watermarking and detection framework that offers strong statistical guarantees. Specifically, we introduce watermarking schemes and hypothesis tests that control the family-wise error rate across multiple reviews, achieving higher statistical power than standard corrections such as Bonferroni, while making no assumptions about the nature of human-written reviews. We explore multiple indirect prompt injection strategies -- including font-based embedding and obfuscated prompts -- and evaluate their effectiveness under various reviewer defense scenarios. Our experiments find high success rates in watermark embedding across various LLMs. We also empirically find that our approach is resilient to common reviewer defenses, and that the bounds on error rates in our statistical tests hold in practice. In contrast, we find that Bonferroni-style corrections are too conservative to be useful in this setting.

大模型检测同行评审水印技术AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。