让AI像专家一样深度思考,提升论文评审质量
DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
- 分阶段模拟专家评审流程,融合文献检索与证据论证
- 140亿参数模型在对比中胜过700亿参数模型,胜率超80%
- 开源完整资源,适合研究者复现与工具开发
大型语言模型(LLMs)在科研评估中应用日益广泛,尤其在自动化论文评审方面。然而,现有基于LLM的评审系统存在领域知识有限、推理幻觉及缺乏结构化评估等问题。为此,我们提出DeepReview,一种多阶段框架,通过结构化分析、文献检索和基于证据的论证来模拟专家评审过程。利用包含结构化标注的DeepReview-13K数据集,我们训练了DeepReviewer-14B模型,在最佳模式下对GPT-o1和DeepSeek-R1的胜率分别达到88.21%和80.20%,且使用更少的计算资源。本工作为基于LLM的论文评审设立了新基准,所有资源均已公开,代码、模型、数据集及演示可访问http://ai-researcher.net。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly utilized in scientific research assessment, particularly in automated paper review. However, existing LLM-based review systems face significant challenges, including limited domain expertise, hallucinated reasoning, and a lack of structured evaluation. To address these limitations, we introduce DeepReview, a multi-stage framework designed to emulate expert reviewers by incorporating structured analysis, literature retrieval, and evidence-based argumentation. Using DeepReview-13K, a curated dataset with structured annotations, we train DeepReviewer-14B, which outperforms CycleReviewer-70B with fewer tokens. In its best mode, DeepReviewer-14B achieves win rates of 88.21\% and 80.20\% against GPT-o1 and DeepSeek-R1 in evaluations. Our work sets a new benchmark for LLM-based paper review, with all resources publicly available. The code, model, dataset and demo have be released in http://ai-researcher.net.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。