提出首个评估人机协作评审的基准,揭示当前检测方法误判根源。
PeerPrism: Peer Evaluation Expertise vs Review-writing AI

- 构建2万+评审数据集,区分思想来源与文本生成来源
- 发现当人类提供观点、AI写文本时,检测器常误判
- 强调需将作者身份视为多维属性,而非二元判断
大型语言模型(LLMs)正广泛用于科学同行评审,协助起草、重写、扩展和润色。然而,现有LLM检测方法大多将作者身份简化为二元问题——人类或AI,未考虑现代评审流程中人机混合的复杂性。实际上,评审观点与文本表达可能来自不同来源,形成连续的人机协作谱系。本文提出PeerPrism,一个包含20,690条同行评审的大规模基准,专门用于解耦思想来源与文本来源。通过设计从纯人类到纯合成再到多种混合生成的受控生成场景,系统评估检测器是否能准确识别表面文本来源或评价推理来源。在PeerPrism上对主流检测方法进行评测发现:尽管部分方法在标准二元任务(人类 vs. 完全合成)中表现良好,但在混合场景下预测结果显著分歧。尤其当观点源自人类而文本由AI生成时,检测器频繁出现矛盾分类。结合风格学与语义分析,结果表明当前方法混淆了文本实现与智力贡献。我们证明,同行评审中的LLM检测不能简化为二元归属问题,而应建模为涵盖语义推理与风格实现的多维构造。PeerPrism是首个评估此类人机协作场景的基准,所有代码、数据、提示与评估脚本已开源,支持可复现研究。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used in scientific peer review, assisting with drafting, rewriting, expansion, and refinement. However, existing peer-review LLM detection methods largely treat authorship as a binary problem-human vs. AI-without accounting for the hybrid nature of modern review workflows. In practice, evaluative ideas and surface realization may originate from different sources, creating a spectrum of human-AI collaboration. In this work, we introduce PeerPrism, a large-scale benchmark of 20,690 peer reviews explicitly designed to disentangle idea provenance from text provenance. We construct controlled generation regimes spanning fully human, fully synthetic, and multiple hybrid transformations. This design enables systematic evaluation of whether detectors identify the origin of the surface text or the origin of the evaluative reasoning. We benchmark state-of-the-art LLM text detection methods on PeerPrism. While several methods achieve high accuracy on the standard binary task (human vs. fully synthetic), their predictions diverge sharply under hybrid regimes. In particular, when ideas originate from humans but the surface text is AI-generated, detectors frequently disagree and produce contradictory classifications. Accompanied by stylometric and semantic analyses, our results show that current detection methods conflate surface realization with intellectual contribution. Overall, we demonstrate that LLM detection in peer review cannot be reduced to a binary attribution problem. Instead, authorship must be modeled as a multidimensional construct spanning semantic reasoning and stylistic realization. PeerPrism is the first benchmark evaluating human-AI collaboration in these settings. We release all code, data, prompts, and evaluation scripts to facilitate reproducible research at https://github.com/Reviewerly-Inc/PeerPrism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。