arXiv:2608.14625cs.CYcs.AI2026-08综述

用本地大模型预审论文,让同行评审更高效透明。

Local AI pre-screening for human triple-blind peer review in health sciences

论文配图:Local AI pre-screening for human triple-blind peer review in health sciences
图 1 · 摘自论文原文
  • 五阶段流程:先去匿名,再多模型本地并行预审,自动校验后交人类盲审。
  • 实测21%的顶会评审已由AI生成,本方案可将平均审稿时间从13周降至更短。
  • 所有AI在本地运行,保障数据安全,适合注重隐私的医学期刊使用。

学术同行评审面临巨大压力:NeurIPS 2025有21,575篇投稿,ICLR 2025为11,603篇,ICML 2025为12,107篇。投稿量已超过合格审稿人供给,大语言模型(LLMs)正悄然填补空缺,但多数未被披露。对ICLR 2026的独立分析显示,约21%的75,800份评审报告为全AI生成,超半数存在AI参与(2024年为15.8%)。风险包括伪造引用及隐藏提示注入指令操控AI评分。本文提出一个三盲、多LLM本地预审框架,专为健康科学期刊设计,明确披露AI参与,同时保留人类审稿人最终决策权。流程包含五个阶段:文档净化/匿名化、并行AI预审、自动校验关卡、盲审人类评审、编辑裁定,并在校验与编辑阶段设置返稿循环。针对美国国立卫生研究院(NIH)/国家科学基金会(NSF)禁止向第三方生成式AI提交未发表提案的保密顾虑,所有三名AI审稿人均运行于本地托管的开源权重大模型,确保稿件内容保留在期刊系统内。此前研究(Shen et al.)在200篇稿件上测试五种开源模型的四分位分类性能,准确率仅35%精确匹配,不支持自主使用,支持本方案保留强制人工裁定。该透明、人类监督的设计为当前隐蔽、无监管的AI评审提供可辩护替代方案,有望在不取代人类判断的前提下,显著缩短传统评审的平均13周首次决定周期。

原文摘要 · Abstract (English)

Academic peer review is under mounting strain: NeurIPS 2025 received 21,575 submissions, ICLR 2025 received 11,603, and ICML 2025 received 12,107. This volume has outpaced the supply of qualified reviewers, and large language models (LLMs) are already filling the gap, largely undisclosed. An independent analysis of ICLR 2026 found roughly 21% of its 75,800 peer reviews were fully AI-generated, with over half showing some AI involvement (up from 15.8% in 2024). Documented risks include hallucinated citations in accepted papers and hidden prompt-injection instructions embedded in manuscripts to manipulate AI reviewers into favorable assessments. We propose a triple-blind, multi-LLM pre-screening framework for peer review, developed for a health sciences journal, that formalizes and discloses AI involvement while preserving human reviewers as the final decision-making authority. The framework routes a submission through five stages -- sanitization/anonymization, parallel AI pre-screening, an automated check gate, blinded human review, and editorial adjudication -- with return-to-author loops at the check and editor stages. Addressing the confidentiality concerns behind NIH/NSF bans on submitting unpublished proposals to third-party generative AI, all three AI reviewers run on locally-hosted, open-weight LLMs, keeping manuscript content within the journal infrastructure. The closest precedent, Shen et al., benchmarked five open-source LLMs on quartile classification of 200 manuscripts and found accuracy insufficient (35% exact-match) for autonomous use, supporting our decision to retain mandatory human adjudication. This transparent, human-supervised design offers a defensible alternative to today's opaque, unregulated AI use in peer review, potentially reducing the substantial delay of traditional review (avg. 13 weeks to first decision) without displacing human judgment.

AI审稿医学期刊本地部署透明评审

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。