构建首个面向学术审稿幻觉的评测基准,助力识别大模型生成的虚假评论。
HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

- 基于真实论文与审稿意见,构建带幻觉注入的三元组数据集。
- 在12,000篇论文、38,000份审稿中发现幻觉普遍存在且难辨识。
- 提出审稿场景专属幻觉分类体系,适合评估和改进审稿AI系统。
学术审稿规模扩大促使使用大语言模型(LLMs)作为审稿助手,但其可能生成流畅却无依据的陈述,损害评审可靠性。现有幻觉评测基准未针对审稿场景设计,因验证需基于长篇技术文献。我们提出HalluPeer,一个用于检测科学审稿中幻觉的基准,提供论文内容、人工撰写的审稿意见与注入幻觉的审稿意见组成的对齐三元组,并进行幻觉检测、分类与定位标注。我们的流程构建了专属于审稿场景的幻觉分类体系,识别审稿上下文并自动过滤后注入幻觉。在12,000篇论文和38,000份审稿上的实验表明,现有检测器难以区分幻觉与合理批评;对真实审稿的评估也证实,HalluPeer定义的幻觉模式确实在实际审稿中出现,凸显源意识验证的必要性。
原文摘要 · Abstract (English)
The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported claims that undermine review reliability. Existing hallucination benchmarks are not designed for peer review, where verification requires grounding claims in long, technical papers. We introduce HalluPeer, a benchmark for detecting hallucinations in scientific peer reviews, providing aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Our pipeline induces a peer-review-specific hallucination taxonomy, identifies review contexts, and injects hallucinations with automated filtering. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, while evaluation on authentic reviews demonstrates that HalluPeer-defined hallucination patterns occur in real peer reviews, highlighting the critical need for source-aware verification. Our project page can be found in https://github.com/Lin-TzuLing/HalluPeer.git
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。