论文提出防御AI代审的隐藏防护机制,让投稿文件自动识破并阻断聊天机器人评审。
Shattering the Echo Chamber: Hidden Safeguards in Manuscripts Against the AI Takeover of Peer Review

- 在PDF文档结构中嵌入不可见文本标记,利用视觉与结构分离特性干扰AI评审。
- 实测在12个领域、7种商用AI下防御成功率最高达84%,人类评审不受影响。
- 适合期刊编委会部署,无需硬件升级,每篇处理仅多耗1秒。
随着大模型能力提升,学术编辑和会议委员会日益担忧审稿人完全委托商业聊天机器人完成同行评审。这一担忧源于先前研究发现,当前聊天机器人缺乏独立批判性思维和深度推理能力来评估科学新颖性。一种有前景的缓解方向是将隐藏指令嵌入稿件,以干扰或改变由聊天机器人生成的评审内容。然而,现有方法仍依赖直观且脆弱的同质化注入策略,通常以跨流方式插入,易被清洗或中和。本文识别出端到端审稿外包这一新兴威胁,提出IntraGuard——一种黑盒、适配所有会议的防御框架,基于PDF固有的结构-视觉解耦特性设计。该框架支持显式(触发拒绝或警告)和隐式(嵌入预定义文本标记)两种策略,可通过三种内流注入机制实现,无缝嵌入异构防御文本对象,不改变文档视觉呈现。在7个真实商用聊天机器人设置和12个跨学科会议上的广泛评估显示,IntraGuard最高防御成功率达84%,同时保持对人类审稿人的评审不变性。该方案轻量且硬件无关,平均仅需1秒/篇即可在普通个人电脑上完成。我们进一步测试了11种自适应攻击,涵盖稿件清洗与指令干扰,并讨论构建集成防御的潜力。
原文摘要 · Abstract (English)
As LLMs become increasingly capable, editorial boards and program committees are growing concerned about reviewers who fully outsource peer review to commercial chatbots. This concern stems from prior findings that current chatbots lack the independent critical thinking and depth of reasoning required to assess scientific novelty. One promising direction for mitigating this concern is to embed hidden instructions into manuscripts that disrupt or alter chatbot-generated reviews. However, existing methods remain intuitive and fragile, as they typically rely on homogeneous payloads injected in an inter-stream manner, rendering them susceptible to sanitization or neutralization. In this paper, we identify End-to-End Review Outsourcing as an emerging threat and propose IntraGuard, a black-box, venue-agnostic defense framework grounded in the structural--visual decoupling inherent to the PDF. Designed for committee-side deployment, IntraGuard supports both explicit strategies that trigger refusal or warning signals, and implicit strategies that embed predefined textual markers into the generated review. These strategies can be deployed via any of three intra-stream injection mechanisms, each of which seamlessly embeds heterogeneous defensive text objects within the PDF's underlying structure without altering its visual presentation. Extensive evaluations across 7 real-world commercial chatbot settings and 12 venues spanning diverse disciplines show that IntraGuard achieves a defense success rate of up to 84%, while preserving peer-review invariance for human reviewers. IntraGuard is lightweight and hardware-independent, incurring an average overhead of only one second per manuscript on a commodity personal computer. We further evaluate 11 adaptive attacks spanning manuscript sanitization and instruction interference, and discuss the implications of constructing ensemble defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。