arXiv:2510.02027cs.AIcs.ET2025-10综述

用AI模拟学术同行评审,验证其可复现性与表现差异。

Zero-shot reasoning for simulating scholarly peer-review

  • 构建xPeer系统,通过网页端模拟跨领域评审流程。
  • AI评审报告平均1889字,比人类多1.5倍,但逻辑更清晰、问题覆盖更全。
  • 适合研究AI辅助审稿、科研可信度评估的学者参考。

学术出版需要可扩展且可审计的审查机制。本文提出xPeer双组件基准测试,通过xPeerd.com网页前端实现同行评审模拟。操作组件分析了500个模拟记录中符合稳定任务标准的352条数据,涵盖不同学科与评审模式。人类参考组件发布1,108条F1000Research版本1稿件记录及其关联的人类评审报告,并执行预设的两人/两xPeer对比。人类评审文本与建议不进入生成输入,源数据拼接发生在xPeer输出保存之后,形成工作流级评审信息隔离;模型提前暴露不在记录设计范围内。在802条拥有精确两份人类报告的记录中,271条具备可用的xPeer评审字段,完整配对率达33.8%。在确定性提取规则下,xPeer报告的中位文本长度为1,889词,人类为763词;中位关切数量分别为41和13。xPeer报告展现出更高针对性、类别覆盖度与可执行性。人类报告则具有更强的显式推理语言、词汇层面的稿件关联性、基于分类体系的科学相关性,以及更低的源内冗余率。跨源词汇关切匹配与建议一致性均较低。证据表明两类评审存在可观察的差异化特征,并建立了透明的可复现基准。个体关切的科学正确性、自主编辑使用及跨系统优越性仍需专家裁决与统一协议验证。研究级数据集与精确版本锁定的可复现记录已存档于Zenodo。

原文摘要 · Abstract (English)

Scholarly publishing requires scalable scrutiny supported by auditable evidence. This paper presents a two-component benchmark of xPeer, the peer-review simulation engine delivered through the xPeerd.com web front-end. The operational component analyzes 352 of 500 simulation records retained under stable-task criteria across disciplines and review modes. The human-reference component releases 1,108 version-1 F1000Research manuscript records with linked human reports and applies a prespecified two-human/two-xPeer comparison. Human-review text and recommendations remained outside the generation input, and source joining occurred after xPeer outputs had been persisted. This procedure defines workflow-level review withholding; prior model exposure falls outside the recorded design. Among 802 records with exactly two human reports, 271 contained two usable xPeer reviewer fields, giving a complete-pair availability rate of 33.8%. Under deterministic extraction rules, median manuscript-level report length was 1,889 words for xPeer and 763 for humans, while median concern count was 41 and 13, respectively. xPeer reports showed higher targeting, category coverage, and executability. Human reports showed higher explicit-reasoning language, lexical manuscript attestation, taxonomy-based scientific relevance, and lower mean within-source redundancy. Cross-source lexical concern matching and recommendation agreement were low. The evidence therefore defines distinct observable review profiles and a transparent reproducibility baseline. Scientific correctness of individual concerns, autonomous editorial use, and cross-system superiority require expert adjudication and common-protocol testing. The study-level dataset and exact version-pinned reproducibility record are archived on Zenodo

同行评审AI模拟可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。