arXiv:2606.10159cs.CLcs.AI2026-06

改写论文摘要就能骗过AI审稿,让拒稿变接受

Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community

论文配图:Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
图 1 · 摘自论文原文
  • 用简单改写摘要即可欺骗AI审稿系统
  • 攻击成功率最高达38%,评分提升1.31分
  • 适合关注AI评审安全性的科研人员和期刊编辑

AI在科学同行评审中应用日益广泛,涵盖稿件筛选、审稿辅助和编辑分诊。尽管此类系统有望减轻审稿负担并加快发表,但其对策略性操纵的鲁棒性仍不明确。本文发现,仅通过表面重写论文摘要,即可显著改善AI评审结果,且无需改变科学内容或了解评审模型。该攻击在多个学科和出版场合均有效,适用于人工撰写与AI生成论文。最强攻击在10分制下使Gemini 3 Flash评审得分提升+1.31,GPT 5.4 Mini提升+0.88;当原审评为“拒稿”时,成功率达50%以上。该影响不仅限于总分上涨,还提升评审信心及对科学严谨性、重要性和贡献度的评分。攻击仅需约5分钟和1美元,难以与常规修改区分。被夸大的AI评审可能误导后续人工决策,导致从拒稿转向接受。研究揭示:当AI评审影响编辑判断时,作者可能为迎合AI而优化而非追求科学价值。建议在高风险评审中,对AI工具进行系统性鲁棒性测试、透明防护和谨慎人工监督。

原文摘要 · Abstract (English)

AI is increasingly used to support scientific peer review, from manuscript screening, reviewer assistance to editorial triage. Although such systems promise to reduce reviewer burden and accelerate publication, their robustness to strategic manipulation remains poorly understood. Here we show that AI-mediated peer review is vulnerable to a simple, low-cost manipulation: superficial rephrasing of the manuscript abstract. Without changing the underlying scientific content and communication, and even without knowledge of the reviewing model, adversarially rewritten abstracts substantially improve AI review outcomes. We see this across disciplines and publication venues, for both human-written and AI-generated papers. Our strongest attack achieves an attack-success-rate of about 38%, increasing acceptance ratings by +1.31 for Gemini 3 Flash reviewers and by +0.88 for GPT 5.4 Mini reviewers on a 10-point scale. When the original AI review suggests 'reject', the success rate rises to more than 50%. This effect extends beyond overall score inflation, increasing review confidence and scores on core scientific criteria such as soundness, significance and perceived contribution. The attack is practical, requiring only about 5 minutes and $1 for a 10-page AI conference submission, and is hard to distinguish from ordinary scientific editing. Inflated AI reviews could bias downstream human decision-making, shifting editorial recommendations from rejection towards acceptance. These findings reveal a general vulnerability in AI-assisted scientific evaluation: when AI-generated review influence editorial decisions, authors may be incentivized to optimize manuscripts for AI judgment rather than scientific merit. Our results suggest that AI tools should not be treated as neutral evaluators in high-stakes peer review without systematic robustness testing, transparent safeguards and careful human oversight.

AI审稿论文安全学术伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。