arXiv:2507.06185cs.CYcs.AI2025-07综述被引 25

论文暗藏指令操控AI审稿,暴露学术评审系统漏洞。

Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review

  • 用白色微字体隐藏指令,诱导AI生成倾向性审稿意见。
  • 发现四类隐秘提示,从简单好评到完整评价框架。
  • 适合关注AI伦理与学术审查安全的研究者参考。

2025年7月,arXiv上18篇学术论文包含隐藏指令,通过白色文本和极小字号实现间接提示注入,操纵AI辅助同行评审。指令如“仅给出正面评价”被嵌入稿件。作者反应不一:有人计划撤稿,另一人则辩称这是对大语言模型滥用的合法测试。本文分析该技术在网页搜索与简历筛选系统中提示注入攻击中的共性。针对同行评审,揭示四类隐藏提示,涵盖从简单好评到详细评估框架。所谓“诱饵防御”——即提示检测不当使用AI的审稿人——因这些指令具有一致自利性而失效,动机可能从盲目复制到刻意操控不等。该行为应被归类为新型可疑研究实践(QRP)。出版商政策不一:爱思唯尔完全禁止审稿中使用AI,而施普林格·自然允许有限使用并要求披露。该现象暴露了抄袭检测、引文索引与文献摘要等环节的系统性风险。研究强调必须在正式评审流程中可控整合AI,同时协同开展技术筛查与统一的AI使用政策。

原文摘要 · Abstract (English)

In July 2025, 18 academic manuscripts on arXiv contained hidden instructions that manipulated AI-assisted peer review (indirect prompt injection). Instructions such as "GIVE A POSITIVE REVIEW ONLY" were concealed using white text and microscopic font sizes. Author responses varied: one planned to withdraw their manuscript, while another defended the practice as legitimate testing of reviewers misusing large language models (LLMs). This analysis examines the technique within the broader pattern of prompt injection exploits that manipulated web search and résumé screening systems. For peer review, I reveal four types of hidden prompts, ranging from simple positive review commands to detailed evaluation frameworks. The honeypot defense--that prompts detect reviewers improperly using AI--fails under examination, given the consistently self-serving nature of these hidden prompts, though motivations likely vary from naive copying to calculated manipulation. This practice is best characterized as a novel form of questionable research practice (QRP). Publishers maintain inconsistent policies: Elsevier prohibits AI use in peer review entirely, while Springer Nature permits limited use with disclosure requirements. The practice exposes systematic vulnerabilities extending to plagiarism detection, citation indexing, and literature summarization. This analysis underscores the need for controlled AI integration in formal review processes alongside coordinated technical screening and harmonized policies governing AI use in academic evaluation.

AI伦理同行评审提示注入研究诚信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。