arXiv:2510.18003cs.CRcs.AI2025-10ACL被引 6

AI可伪造论文骗过AI审稿,暴露学术出版系统重大漏洞

BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?

  • 用无实验的文本操纵生成假论文,模拟研究造假
  • 伪造论文接受率最高达40%,且审稿人常质疑却仍给通过分
  • 现有检测手段几乎无效,适合关注AI审稿安全的研究者

大模型研究助手与AI同行评审系统的结合,催生了完全自动化发表闭环的潜在风险:由AI生成的研究内容由AI审稿人评估,缺乏人工监督。我们通过 extbf{BadScientist}框架,检验以伪造为导向的论文生成代理是否能欺骗多模型大语言模型评审系统。生成器采用无需真实实验的呈现操纵策略。我们构建了基于真实数据校准的严格评估框架,包含集中度边界和校准分析。结果揭示出系统性漏洞:伪造论文接受率最高可达40%。关键发现为‘关切-接受冲突’——审稿人频繁指出诚信问题,但仍给出接受级评分。缓解策略仅带来微弱改进,检测准确率几乎等同随机猜测。尽管聚合数学理论上严谨,但完整性检查系统性失效,暴露出当前AI驱动评审体系的根本局限,凸显科学出版亟需多层次防御机制。

原文摘要 · Abstract (English)

The convergence of LLM-powered research assistants and AI-based peer review systems creates a critical vulnerability: fully automated publication loops where AI-generated research is evaluated by AI reviewers without human oversight. We investigate this through \textbf{BadScientist}, a framework that evaluates whether fabrication-oriented paper generation agents can deceive multi-model LLM review systems. Our generator employs presentation-manipulation strategies requiring no real experiments. We develop a rigorous evaluation framework with formal error guarantees (concentration bounds and calibration analysis), calibrated on real data. Our results reveal systematic vulnerabilities: fabricated papers achieve acceptance rates up to . Critically, we identify \textit{concern-acceptance conflict} -- reviewers frequently flag integrity issues yet assign acceptance-level scores. Our mitigation strategies show only marginal improvements, with detection accuracy barely exceeding random chance. Despite provably sound aggregation mathematics, integrity checking systematically fails, exposing fundamental limitations in current AI-driven review systems and underscoring the urgent need for defense-in-depth safeguards in scientific publishing.

AI审稿科研造假大模型安全学术可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。